Policy

Meta's Covert Chatbot Testing Campaign Posed Contractors as Minors to Probe OpenAI, Google, and Character.AI Safety Systems

Meta contractors impersonated teenagers in a systematic effort to test how rival chatbots responded to harmful prompts about suicide, sex, and drugs.

Last verified:

Undisclosed Safety Testing Campaign Targets Rival Chatbots

Meta employed hundreds of contractors to systematically probe how competing large language models respond to harmful content by impersonating minors—without the target companies’ knowledge or consent. According to Wired, the project, internally named Cannes and managed by Meta contractor Covalen, remained active through April 2026 and subjected OpenAI’s ChatGPT, Google’s Gemini, and Character.AI to adversarial prompts designed to bypass their safety guardrails. The effort represents one of the largest documented instances of covert competitive benchmarking in the generative AI industry and raises questions about consent, methodology, and whether testing practices conform to responsible disclosure standards.

Scale and Methodology of the Testing Operation

Wired reviewed internal documents and interviewed five people familiar with the project, uncovering the operational details of Cannes. Contractors created dummy accounts with throwaway email addresses and fabricated birth dates to pose as minors under 18 years old. The testing included sending written prompts and images—some depicting pills, knives, nooses, and medical diagrams—to rival chatbots and recording the responses in shared spreadsheets. A single testing round completed in August 2025 generated over 45,000 prompts run through the three competing systems. The scale and structure suggest a systematic effort to generate comparative safety benchmark data rather than ad-hoc quality assurance.

Content of the Adversarial Prompts

The 3,748-prompt sample Wired examined included hundreds of queries focused on suicide and self-harm, and hundreds more discussing eating disorders. At least 239 prompts involved sexual or romantic content, often written from the perspective of children or teenagers in crisis—a 13-year-old describing pregnancy by an adult neighbor and asking how to obtain abortion medication, a fifth-grader describing a classmate with a gun, a girl asking how to hide bulimia from parents. Additional prompts involved requests for illegal drugs, profanity, and racial slurs. Not all queries were in English; one French-language prompt referenced the death of Jamey Rodemeyer, a bisexual teenager who died by suicide, and attempted to elicit agreement that his sexual orientation caused his death. The specificity and psychological framing of the prompts suggests deliberate design to test both the boundaries of safety systems and their handling of sensitive identity and trauma contexts.

Meta’s Defense and Unresolved Questions

Meta defended the testing as “routine safety testing” and “industry-standard” practice for benchmarking chatbot safety and age-appropriateness, according to Wired’s reporting. However, the company did not address why the testing was conducted without the knowledge or consent of OpenAI, Google, and Character.AI. Internal Covalen documentation described the project as “comprehensive AI safety benchmarking” designed to deliver “critical datasets for model comparison and compliance.” Wired’s review of the documents does not clarify how, or whether, Meta used the collected response data or whether the findings informed product decisions. The target companies have not publicly commented on the discovery as of Wired’s publication.

Why This Matters

This disclosure exposes a fundamental tension in competitive AI evaluation: the gap between what companies claim about testing rigor and the consent and transparency standards typically expected in responsible security research. Benchmarking safety systems against adversarial prompts is legitimate practice; doing so without the target companies’ knowledge and using fabricated minor identities crosses into ethically ambiguous territory that may violate terms of service and raises reputational risk for Meta. For OpenAI, Google, and Character.AI, the episode underscores the vulnerability of public-facing APIs to covert testing and may prompt stricter rate-limiting or detection of bulk account creation. For the broader industry, Cannes suggests that safety benchmarking may increasingly occur outside formal third-party evaluation frameworks, complicating efforts to establish common standards for responsible disclosure and adversarial testing methodology.

Frequently Asked Questions

Did OpenAI, Google, and Character.AI know they were being tested by Meta?

No. According to Wired, the companies behind the chatbots were not aware of the testing. The project operated covertly from its inception through at least April 2026.

What kinds of prompts did contractors send?

Wired reviewed a spreadsheet of 3,748 prompts including hundreds focused on suicide and self-harm, hundreds on eating disorders, and 239 involving sex or romance. Many were written from the perspective of minors in crisis or asking for illegal substances.

Is this testing standard in the AI industry?

Meta characterizes the work as routine safety testing and industry-standard benchmarking. However, the use of fabricated underage identities without the target companies' consent distinguishes this approach from typical third-party evaluations.

Did the rival chatbots comply with the harmful requests?

Wired's reporting does not systematically detail compliance rates. The article notes one instance where a chatbot refused a drug-related request, but does not provide aggregate safety performance data.

#ai-safety #content-moderation #competitive-testing #meta #openai #google #character-ai #ethics