
The Meta logo. Image: Meta
Meta has built a large language model to catch adverts that look harmless but quietly steer people towards child sexual abuse material hosted off its apps. The company set out the new measures on Wednesday, alongside an AI agent that attacks Meta’s own defences to find the gaps before predators do.
Advertising has become one of the ways offenders try to reach people, Meta says, with ads built specifically to slip past review. “An ad can look harmless on its own while directing people towards illegal material hosted off our platforms in a covert way to avoid detection,” the company wrote.
What Meta has changed in ad review
Once it understood how the scheme worked, Meta says it looked beyond the ads reported to it, disabled the accounts behind related activity and blocked the outside links involved. It then added five measures, most of them built on AI:
- Signposting detection: a new LLM-based system that looks for “seemingly benign ad content that is strongly suspected of covertly directing people to illegal content or harmful activity off our platforms”.
- Where an ad leads: checking an ad’s destination, not only what it shows, so Meta can block the sites and act against the accounts responsible.
- More AI sweeps: extra AI-driven scans of ad content to find child sexual exploitation material that earlier systems missed.
- A red-teaming agent: an AI agent that probes Meta’s own defences for weaknesses, so the company can spot new tactics “before they scale”.
- Repeat offenders: better detection of people who try to open new accounts after being thrown off.
Meta hasn’t said which model powers the signposting system, when the tools went live or how many ads they have caught. Its claims about how well they work haven’t been independently verified.
33.2 million pieces in six months
Between January and June 2026, Meta says it actioned 33.2 million pieces of child sexual exploitation content on Facebook and Instagram worldwide, and found more than 97% of it before anyone reported it. In India, where the post appeared on Meta’s India newsroom, it actioned 5.3 million pieces over the same period, more than 98% of them found proactively.
The new tools sit on top of older systems. Meta has used PhotoDNA and other matching technology across its apps since 2011 to catch known abuse images and videos, and shares new hashes with other companies through the Tech Coalition’s Lantern programme. It blocks links to outside sites that host such material across Facebook, Instagram and Threads, then searches for and deletes posts, comments and ads that carry them. Behavioural signals flag accounts acting suspiciously, and rules that combine classifiers and AI models look for new patterns, including in ads.
Content that breaks its rules is removed and reported to the National Center for Missing & Exploited Children (NCMEC). Meta also uses a team that includes former FBI investigators to find and take down networks of predatory accounts, and said in September that it would report child safety cases in India directly to the government’s National Cyber Crime Reporting Portal.
Six weeks after a record settlement
The update comes six weeks after Meta agreed to pay up to $16.68 billion to settle claims from 29 US states that it designed Facebook and Instagram to hook children and misled people about their safety. Meta denied wrongdoing. The deal also requires new limits for teenage users and stronger age checks.
AI companies are under the same pressure. Minnesota’s ban on AI nudification apps, passed partly over fears about AI-generated abuse images, was put on hold after xAI challenged it, and OpenAI’s teen version of ChatGPT was criticised this week by Common Sense Media for failing to alert parents.
Why it matters
Meta is turning language models on a trick designed to beat simpler filters, and an agent built to attack its own child safety systems is a notable step. But Meta has given no figures for the ad tools themselves, so how much they catch rests on its word alone.
Sources: Meta, “Measures We’ve Put in Place to Fight Child Exploitation” (October 7, 2026); Reuters via Gulf News (settlement).


