Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Inferential Privacy Leakage in Anonymized Conversational AI Logs

Domain:

natural language processingdigital infrastructure

Record type:

paperdataset
Creator:
ZamGar
Host:avatar
Hundreds of millions of users now hold detailed, multi-turn conversations with ChatGPT and similar LLM assistants. We measure two privacy-relevant features of these conversations on a corpus of complete ChatGPT histories donated by over 1,000 users in four Global South countries (Brazil, India, Nigeria, Pakistan). First, on explicit disclosure: 34.5% of user messages contain personal information across a twenty-category taxonomy, with the median user first revealing identifying content within the first 14% of their conversation history. Second, on inference beyond explicit disclosure: we restrict to a cohort whose conversations contain no messages flagged by an LLM-based filter for explicit demographic self-identification (a separate NER pass marks PII for the disclosure audit but does not drive cohort exclusion). On this filtered cohort, an off the shelf large language model still recovers each user's age, gender, and country at weighted F1 of 0.84, 0.90, and 0.88, respectively, with the median user identified from the first 5% of their conversation history. Reading the model's natural-language reasoning traces, we identify four recurring stereotype patterns that drive both successful inference and an asymmetric error distribution concentrating on women in technical fields, older users with contemporary skills, and Global South tech professionals. We also compare ChatGPT against the same users' Google Search and YouTube histories as inference surfaces, and find it competitive with these older substrates that have driven behavioral advertising for two decades. Message-level PII removal is insufficient on its own as a privacy intervention for conversational AI data.

Visit

arxiv.org

Tasks

information extractionnamed entity recognition

Tags

Computers and SocietySocial and Information Networks

Similar

Inferential and counter-inferential grammatical markers in Swahili dialogueOpenray-ai/naija-privacy-filterInsights into Privacy Protection Research in AIBytte-AI/Pidgin-to-English-conversational-translationskefasmanu/Multi-lingual-Conversational-AI-Chatbot-ACEDSThe Inferential Evidential in Sahidic Coptic

Inferential and counter-inferential grammatical markers in Swahili dialogue

Includes bibliographical references (leaves 14-15) Photocopy of manuscript submitted for publication

Openray-ai/naija-privacy-filter

LoRA adapter for openai/privacy-filter on Nigerian-domain PII detection. # Naija Privacy Filter `n

Insights into Privacy Protection Research in AI

This paper presents a systematic bibliometric analysis of the artificial intelligence (

Bytte-AI/Pidgin-to-English-conversational-translations

Sample dataset: Nigerian Pidgin to English translation pairs for machine translation research 🤗 Hugg

kefasmanu/Multi-lingual-Conversational-AI-Chatbot-ACEDS

This Project is a Dissertation Implementation of a multilingual conversational AI chatbot for enhanc

The Inferential Evidential in Sahidic Coptic

International audience no abstract