HomeNews ReportsHow global tech giants are flagging child abuse material and helping to catch criminals

How global tech giants are flagging child abuse material and helping to catch criminals

India received nearly 1.93 million NCMEC CyberTipline reports in 2025, while global platforms used hashing, artificial intelligence, behavioural signals and human review to identify known and previously unseen abusive material linked to users in India.

It was May 2025. There was no complaint from a child in Aizawl, Mizoram. His family did not approach the police. However, thousands of kilometres away, pieces of a digital trail were already sitting inside systems built to detect the sexual exploitation of children online.

On 30th May 2025, the Central Bureau of Investigation (CBI) registered a case against an unidentified accused over allegations of the creation, collection, possession and dissemination of sexually explicit material involving children. Just four days later, on 4th June, CBI officers searched a residence in Aizawl. They seized the suspect’s electronic devices.

The investigation did not remain limited to finding prohibited images and videos on those devices. According to the CBI, forensic examination revealed material that indicated that a minor had also been sexually assaulted. The agency stated that the material was corroborated using INTERPOL’s International Child Sexual Exploitation database and CyberTipline reports submitted by Google to the Indian Cyber Crime Coordination Centre, or I4C.

The digital trail eventually helped the investigators identify and locate the child. The CBI arrested one accused in Aizawl on 9th June. More importantly, a child who had apparently never approached the police was identified and rescued from further harm.

The Aizawl case needs a deeper discussion, as it became a criminal investigation involving a search, seized devices, forensic examination and an arrest. However, before moving forward, it is essential to understand how material circulating through the internet became visible to a technology company’s child-safety system in the first place.

The answer to this question lies in an international detection infrastructure that operates largely outside the view of an ordinary internet user.

An abusive image uploaded through an account in India may already have been identified somewhere else in the world. A technology company may be able to recognise another copy through a mathematical fingerprint of the image. If the material has never been identified before, artificial intelligence and other automated classifiers may flag it for specialised review. A user may report it. In other circumstances, suspicious activity involving an account or network can provide the first signal.

Where applicable, the resulting information can enter reporting systems such as the CyberTipline operated by the US-based National Center for Missing & Exploited Children, or NCMEC. India has formally received NCMEC CyberTipline reports since the National Crime Records Bureau signed an agreement with the organisation on 26th April 2019.

And the volume is enormous. In 2025, NCMEC associated approximately 1.93 million CyberTipline reports with India. But that figure does not mean 1.93 million children were abused, 1.93 million offenders were identified, or 1.93 million criminal cases existed. NCMEC itself warns that CyberTipline reports and reported files measure reporting activity and cannot be used as a proxy for the number of victims, crimes or police investigations.

It is important to understand the distinction because abusive material has a characteristic that makes the numbers particularly difficult to understand that the material can survive online long after the original abuse is over.

An image can be copied. A video can be redistributed. The same material can appear through different accounts, on different services and at different points in time. One victim’s material can therefore produce thousands, or even far more, individual detections without representing thousands of different victims. In simple words, a single incident can turn into thousands of tipline reports but, in the end, it must be seen as a single incident.

Meta showed the scale of this duplication when it examined child exploitative material it reported during October and November 2020. More than 90% was the same as, or visually similar to, previously reported material. Copies of only six videos accounted for more than half of the material in that sample.

That is why some of the most important technology used against child sexual abuse material is not designed to “look” at every image in the way a human reviewer would. Instead, it can recognise the digital fingerprints of material that has already been identified.

To understand how an online signal involving India can eventually become useful to Indian law enforcement, the first step is therefore to understand what child sexual abuse material is, what technology companies are actually detecting and what their systems can, and cannot, establish.

What exactly are CSAM and CSEAM?

CSAM, or child sexual abuse material, can be defined as an image or video depicting the sexual abuse of a child. It is not merely another category of illegal pornography. CSAM records real abuse and the material potentially documents an offence against a child. Its continued circulation can also perpetuate the exploitation long after the original recording was created.

While CSAM is a widely used term internationally, in India, the Supreme Court emphasised the use of the term CSEAM, or child sexual exploitation and abuse material. In its judgment in Just Rights for Children Alliance vs S Harish on 23rd September 2024, the apex court rejected the expression “child pornography” as inadequate and endorsed the use of CSEAM.

The court said the older terminology could trivialise the seriousness of the offence by associating such material with pornography involving consenting adults.

How CSEAM is detected in India-linked cases and the role of tech platforms

An image or video involving the sexual abuse of a child in India does not necessarily have to be discovered first by an Indian police officer. It may be detected when it appears on the systems of a global technology company, matched with material already identified elsewhere in the world, flagged by an automated classifier or reported by another user. Information connected to that detection can then travel through an international reporting system before reaching Indian authorities.

The scale of that India-facing system is substantial. In 2025, the National Center for Missing & Exploited Children (NCMEC) recorded 19,33,900 CyberTipline reports against India as the relevant country, comprising 17,69,595 referrals and 164,305 reports classified as informational. Among individual countries listed by NCMEC, only the United States had a higher total. However, these figures are not a count of Indian victims or offenders, and NCMEC cautions that geographic attribution can be affected by proxies, anonymisers and other technical factors. A report can also have connections with more than one country.

India has had a formal connection with this system since 2019. On 26th April that year, the National Crime Records Bureau (NCRB) signed an agreement with NCMEC to receive CyberTipline reports concerning online child sexual abuse material and exploitation linked to India. The Union government said that, as of 31st March 2024, more than 69.05 lakh CyberTipline reports had been shared with the concerned States and Union Territories under this mechanism.

This is where the technology behind the system becomes important. An abusive file uploaded from India may be one that has already appeared thousands of times elsewhere. It may instead be material that a platform has never encountered before. In other cases, suspicious behaviour surrounding an account may provide the first signal. Each requires a different form of detection.

One victim’s material can circulate for years

Once an image or video enters circulation, copies can be uploaded repeatedly across accounts and platforms. The same material can consequently generate numerous detections and reports even though it concerns the same underlying act of abuse.

Meta showed this duplication problem through an analysis of material it reported to NCMEC during October and November 2020. The company found that more than 90% of the child exploitative content was the same as, or visually similar to, previously reported content. Copies of only six videos accounted for more than half of the material it reported during that period.

NCMEC provides an even more striking example. In its 2025 CyberTipline report, it said material depicting one child victim had been circulating for around 20 years and had appeared more than 1.4 million times in submissions to NCMEC.

This explains both the enormous numbers generated by online detection systems and the importance of recognising previously identified material automatically. It would make little sense to require a human reviewer to repeatedly rediscover the same abusive image every time another copy surfaced.

It also explains why figures describing “reports”, “files”, “accounts”, “URLs” or “content removed” should never casually be converted into the number of victims or offenders.

There is no single scanner watching the internet

There is no universal CSAM detector through which every file uploaded by an Indian internet user passes. Detection depends on the service being used and its technical architecture. Material can appear on social-media platforms, video-hosting services, cloud-storage products, search engines and other consumer internet services. Users can report content themselves, while companies can also use automated safety technologies.

For an India-linked case, detection can broadly begin through three routes.

The first is known material: imagery that has previously been identified and is appearing again.

The second is previously unknown material: content that is not already represented in a known-CSAM database but which automated systems identify as potentially abusive.

The third involves account, behavioural or network signals: activity suggesting potentially exploitative conduct even where a known-image match was not the original trigger.

The systems can also feed into one another. New material identified today can subsequently become known material, allowing future copies appearing in India or elsewhere to be detected automatically.

How hashing recognises material already identified elsewhere

The easiest way to understand hash matching is to think of a digital fingerprint. When material has already been identified as CSAM, technology can produce a mathematical representation of it known as a hash. A participating service can then compare the fingerprint of material encountered on its systems against fingerprints of known abusive material.

This means, for example, that an image uploaded through an account associated with India need not previously have been investigated by an Indian police force. If the underlying imagery has already been identified elsewhere and its hash is available to a platform, the service may be able to recognise the reappearance.

One of the best-known technologies used for this purpose is PhotoDNA. Microsoft developed it with Dartmouth College in 2009. PhotoDNA creates a digital signature from an image and compares it with signatures of previously identified illegal child sexual exploitation imagery.

Importantly, PhotoDNA is not facial-recognition software. Microsoft explicitly says it cannot identify a person or an object in an image. Nor can a PhotoDNA hash be reversed to recreate the original image.

That is an important distinction in understanding what technology can prove. A PhotoDNA match can help establish that imagery corresponds to known material. It cannot tell police who uploaded it, identify the child simply from the hash or establish who was operating an account at a particular time.

Google uses the same broad principle. It says previously identified CSAM can be automatically flagged through hash matching, while its CSAI Match technology allows participating organisations to identify re-uploads of known abusive video material.

But the obvious question is: who decides what enters a known-CSAM database in the first place?

Google says it obtains hashes from trusted sources, including NCMEC and the Internet Watch Foundation. It also independently reviews purported CSAM hashes before confirmed material is incorporated into its own detection systems. NCMEC separately operates a hash-sharing programme through which participating companies can access fingerprints of identified material.

By the end of 2025, NCMEC said it was sharing more than 12.1 million hashes with 78 electronic service providers participating voluntarily in its hash-sharing initiative. Google, meanwhile, reported that it had cumulatively contributed more than 3.58 million CSAM hashes to NCMEC’s database. This produces an international detection network. Material first identified in one jurisdiction can eventually become recognisable when another copy appears on a participating service being used from India.

How CSEAM is detected.

What happens when the material has never been seen before?

Hash matching has a basic limitation. It cannot find an image by comparing it with a known fingerprint if nobody has previously identified that image. Technology companies therefore use other automated systems to look for potentially new material.

Google says it combines hash matching with artificial intelligence. Its AI systems can flag previously unseen material displaying patterns similar to confirmed CSAM. Google says each new image flagged through this process is reviewed by trained personnel before it is reported as CSAM.

The conceptual difference is straightforward. Hash matching asks: “Does this correspond to something already known?” An automated classifier asks something closer to: “Does this previously unseen material have characteristics suggesting that it falls within the prohibited category?”

Meta has similarly said that it uses artificial intelligence and machine learning, alongside matching technology, to find child-exploitation material and detect potentially inappropriate interactions and other signals involving children.

For an Indian investigation, however, an automated finding is a starting signal rather than the end of the evidentiary process.

A machine flag cannot establish beyond dispute who created a file, who knowingly stored it, who uploaded it or whether the registered owner of an account was the person using it at that time. Those are separate questions which may require investigation, subscriber information, device examination and digital forensics.

In short, machine detection is not a criminal conviction.

Human reviewers remain part of the system

The enormous volume of material makes automation indispensable, but the major platforms do not describe the process as being completely automated.

Google says its detection technologies are supported by specialised content reviewers and subject-matter experts with backgrounds including law, child safety, advocacy, social work and cyber investigations. The company also says counselling, specialist resources and other measures are provided because reviewing potentially abusive material carries an obvious psychological burden.

Human review also provides an important safeguard against incorrect enforcement. Microsoft’s 2025 transparency figures are particularly useful here. The company said it actioned 22,651 consumer accounts associated with child sexual exploitation and abuse imagery, including grooming, across hosted consumer services during the year. Yet 15.08% of accounts actioned were reinstated following appeal and further review.

That does not make automated detection pointless. It demonstrates why platform enforcement, and criminal proof cannot be treated as the same thing. A company can disable an account because its systems and reviewers conclude that platform rules have been violated. That decision may later be reversed on appeal. A criminal prosecution in India must independently satisfy the requirements of Indian law.

Google similarly provides an appeal mechanism for users whose accounts are disabled following CSAM enforcement and says a member of its child-safety team reviews such challenges.

Detection is not limited to looking at images

Some of the more sophisticated child-safety systems do not begin with the contents of an individual file at all. Platforms can examine signals surrounding accounts and networks. These can include reports from users, links associated with prohibited material, repeated policy violations or activity connecting accounts exhibiting suspicious behaviour.

Meta said in July 2026 that it uses technology to identify accounts showing potentially suspicious activity relating to children. It also said that its AI systems can detect suspicious off-platform links when those links appear alongside other signals of child exploitative activity.

This has produced India-specific enforcement. Meta said that, during the six months preceding its July 2026 disclosure, its systems had removed 160,000 accounts in India after advanced AI tools identified suspicious off-platform links in combination with other indicators of child exploitative activity. Those numbers are Meta’s own platform-enforcement figures and do not mean 160,000 Indian offenders were identified or 160,000 criminal cases were registered. They describe accounts removed under the company’s systems and policies.

The detection landscape can therefore be understood as three layers:

Content recognition: Is this known material?

Content classification: Does previously unseen material appear potentially abusive?

Behaviour and network analysis: Is surrounding account activity displaying signals associated with exploitation?

The detection layers.

From a platform somewhere in the world to an Indian reporting system

This is the point at which a global technology system begins to intersect with Indian institutions.

Under US law, covered electronic service providers are required to report specified suspected forms of child sexual exploitation they become aware of to NCMEC’s CyberTipline. NCMEC also accepts reports from members of the public, while non-US companies can voluntarily register to make reports.

NCMEC reviews incoming reports and attempts to determine the location relevant to the incident so that the information can be made available to an appropriate law-enforcement agency. It describes its CyberTipline as effectively serving as a global clearing house because US technology companies have users around the world.

A CyberTip can contain more than an alert saying that a prohibited image was found. Google says a report can, depending on the circumstances, contain information relating to the responsible user, the child, the content and other contextual facts. NCMEC says a report classified as a referral usually contains enough information for potential law-enforcement action, such as user information, imagery and a possible location for a child or offender.

For India, the formal bridge was created through the NCRB-NCMEC agreement signed on 26th April 2019. The Indian Cyber Crime Coordination Centre also states that the arrangement was created to receive CyberTipline reports concerning CSAM related to India.

Meta says that when it becomes aware of apparent child exploitation, it reports through NCMEC in compliance with applicable law. For India specifically, the company says it ensures that such material is reported by NCMEC to India’s National Cyber Crime Reporting Portal on Meta’s behalf.

India also has its own citizen-facing reporting mechanism. The Cyber Crime Reporting Portal was originally launched in September 2018 as a centralised facility for reporting child sexual abuse material and rape or gang-rape related content. The revamped National Cyber Crime Reporting Portal was launched in August 2019 and expanded the system to other categories of cybercrime, with a continued special focus on offences against women and children.

What happens after such information reaches an Indian agency, which State or UT receives it, when verification begins and what turns a digital lead into an FIR are separate questions. Those form the subject of the next part of this series.

India is one of the largest destinations for CyberTipline reports

The global numbers help place India’s 1.93 million reports in perspective.

NCMEC received 21.3 million CyberTipline reports in 2025. Electronic service provider reports contained 61.8 million images, videos and other files, including 29.4 million images and 26.3 million videos. NCMEC said 77% of CyberTipline reports that year involved uploads of CSAM or other child sexual exploitation by users outside the United States.

India alone was associated with 1,933,900 reports in NCMEC’s 2025 country table. Of these, 1,769,595 were categorised as referrals, while 164,305 were informational reports.

NCMEC explains the distinction this way. A referral contains sufficient information for potential law-enforcement action, while an informational report may contain insufficient information or involve material that has become viral and has already been repeatedly reported.

NCMEC explicitly warns against treating these figures as a measure of the prevalence of child sexual abuse in a country. They are not the number of victims, not the number of offenders, not the number of investigations opened and not a comprehensive measurement of the amount of CSAM circulating in a jurisdiction.

That warning is especially important when reporting India’s high numbers. A large CyberTipline count can reflect platform usage, reporting practices, repeat circulation of the same material and the ability to determine a geographic nexus. It cannot simply be presented as “1.93 million cases of child abuse in India”.

What Meta, Google and Microsoft are finding at scale

Meta said that globally it removed 36 million pieces of child-exploitation content from Facebook and Instagram during 2025 and automatically removed more than four million suspicious accounts. During October to December 2025 alone, it removed 13 million pieces of child sexual exploitation content, more than 96% of which it said was detected proactively before anybody reported it.

India is not merely present somewhere within those global figures. Meta’s July 2026 disclosure specifically highlighted the removal of 160,000 India-linked accounts in the preceding six months through systems looking at suspicious off-platform links in combination with other exploitation signals.

Google’s transparency data for July to December 2025 show 862,284 CyberTipline reports to NCMEC and enforcement against 365,597 accounts globally for CSAM violations. Of the content reported to NCMEC during that period, 97.68% was detected through automated systems. India appears first in Google’s displayed list of the top ten countries associated with CSAM-related account enforcement during that six-month period, although the company does not publish the individual India count on that dashboard.

Google also reported and blocked 264,371 CSAM-related URLs from appearing in Search during the second half of 2025. It stresses that blocking a result from Google Search is not the same as removing the underlying material from the third-party website hosting it.

India inside the global detection system

Microsoft reported 111,931 CyberTipline reports to NCMEC during 2025. Across hosted consumer services including OneDrive, Outlook, Skype and Xbox, it actioned 237,391 pieces of content and 22,651 consumer accounts associated with child sexual exploitation and abuse imagery, including grooming. Microsoft said 99.73% of that actioned content was detected through automated technologies.

These figures should not be directly compared as though Meta, Google, Microsoft and NCMEC were measuring the same thing.

A CyberTipline report is not the same as a piece of content. A piece of content is not the same as an account. An account is not necessarily an offender. And none of these units automatically equals a victim or criminal case.

Not every platform can detect the same thing

Another misconception is that because Meta, Google or another large technology company operates in India, it can automatically inspect everything communicated through every one of its services.

Different products have different architectures. Publicly hosted social-media content, a file stored in a cloud account, a search result and private communications are not technically identical environments. What a provider can detect depends on what information is available to it and how the relevant service is built.

This becomes particularly important when discussing encrypted communications. The existence of child-safety systems on one service should not be used to imply that a company necessarily has identical technical visibility across every product it owns.

The same caution applies to the capabilities of individual detection technologies. PhotoDNA does not identify offenders. Hashes do not prove who was holding a device. Classifiers do not establish intent. An account-registration name does not by itself establish who performed an upload.

Microsoft expressly states that PhotoDNA cannot identify people or objects. Google likewise describes its automated systems as mechanisms for detecting and flagging content, supported by human review and subsequent reporting processes.

Once criminal responsibility becomes the question, investigators may need to establish who controlled the account, who possessed the device, what data can be recovered, where the upload originated and whether the necessary ingredients of an offence under Indian law can be established.

That boundary is critical because a technology company detects and reports. A court determines guilt on evidence.

Generative AI is creating a new problem for an old detection system

The detection ecosystem was built largely around identifying and limiting the redistribution of material depicting or derived from child sexual abuse. Generative artificial intelligence has complicated that model.

Google says its systems are designed to detect child sexual abuse and exploitation material including AI-generated CSAM. Its approach includes hash matching, AI classifiers and human review, while its generative-AI safety systems also attempt to block requests seeking the creation of exploitative material.

NCMEC’s 2025 figures show why the issue can no longer be treated as theoretical. It received more than 400,000 reports involving a generative-AI nexus that year, including more than 182,000 reports involving people possessing, generating or attempting to generate AI-related CSAM.

But AI-generated material raises separate questions about synthetic depictions, manipulated images of real children, Indian legal definitions, evidentiary issues and the obligations of AI companies. Those questions go far beyond how traditional hash-matching systems work and will be examined separately in the fourth part of this series.

Technology can identify the signal, but it cannot answer the case

The extraordinary scale of the detection infrastructure means an India-linked file can potentially be recognised without an Indian police officer having previously seen it.

A hash can reveal that imagery corresponds to material already identified elsewhere. AI can surface potentially new material. Human reviewers can assess what machines flag. Behavioural systems can identify suspicious accounts and networks. NCMEC can receive a platform’s report and determine that the information has a geographic connection with India.

India’s own institutional link to that global system has existed formally since the NCRB-NCMEC agreement of 2019. In 2025 alone, NCMEC’s data associated nearly 1.93 million reports with India.

But none of these systems can conclusively answer the questions that ultimately matter in a criminal case.

  • Who produced the material?
  • Who uploaded or circulated it?
  • Who actually controlled the account?
  • Where did the underlying abuse occur?
  • Can the child be identified?
  • Is the child still at risk?
  • What evidence exists on the suspect’s devices?
  • And can Indian investigators establish the offence in court?

Those questions require law enforcement.

A platform detection therefore becomes significant not merely when prohibited material disappears from a screen, but when the digital information surrounding it can help investigators move towards a child who needs protection or a person suspected of committing an offence.

For India, that is where the next stage begins. A CyberTip or other online alert reaches the Indian system, is routed to the relevant police agency, is examined against available digital information and, where sufficient grounds exist, begins the journey towards an FIR, forensic investigation and arrest.

This article is the first part of OpIndia’s 4-article series on CSEAM.

Join OpIndia's official WhatsApp channel

  Support Us  

For likes of 'The Wire' who consider 'nationalism' a bad word, there is never paucity of funds. They have a well-oiled international ecosystem that keeps their business running. We need your support to fight them. Please contribute whatever you can afford

Anurag
Anurag
Anurag is a Chief Sub Editor at OpIndia with over 22 years of professional experience, including more than six years in journalism. He is known for deep dive, research driven reporting on national security, terrorism cases, judiciary and governance, backed by RTIs, court records and on-ground evidence. He also writes hard hitting op-eds that challenge distorted narratives. Beyond investigations, he explores history, fiction and visual storytelling. Email: [email protected]

Related Articles

Trending now

- Advertisement -