
WhatsApp has officially launched 'Scam Alert' in a limited beta rollout on August 12, marking the first published instance on a major encrypted messaging platform where anti-fraud AI and end-to-end encryption operate simultaneously. The optional feature uses an on-device machine learning model that runs directly on users' phones to flag suspected scam messages without ever forwarding message content to Meta's servers. According to the Meta Engineering Blog, the feature downloads a lightweight machine learning model to the device that evaluates incoming messages from non-contacts, looking for conversational structures and linguistic patterns associated with known scam types including urgency cues, requests for personal information, financial solicitations, and extended trust-building rhythms of pig-butchering schemes. As reported by WhatsApp, the model is trained on scam conversations reported by users and is designed to check whether incoming messages from non-contacts match known scam patterns using linguistic signals and probabilistic classification based on conversational structure. With more than 3 billion users, WhatsApp represents one of the largest areas of scam activity, including wire transfers and pig butchering schemes.
The system operates around three main principles: on-device processing, no automatic reporting, and user control. Message content does not leave the device for classification, and the feature does not automatically report messages to WhatsApp, Meta, or any third party. When the model identifies a likely scam, users see a warning banner inside the chat that is visible only to the recipient, never to the sender. Users have four choices: block the contact, report the message to WhatsApp, ignore the warning and continue the conversation, or mark the contact as trusted, which suppresses future alerts for that thread. If users decide that a warning is incorrectly flagged, they can mark the chat as trusted, in which case the warning is removed and Scam Alert will not flag that chat again. According to WhatsApp, if a user marks that they trust a chat, they can opt in to share the last 5 messages received with WhatsApp to help improve the feature's accuracy. The only data that leaves the device is an anonymized, differentially private count of how many warnings were shown and what users did afterward, processed inside a hardware-isolated Trusted Execution Environment before reaching Meta's servers. Cloudflare serves as an independent key holder and ledger operator, meaning even Meta's own engineers cannot certify a model as legitimate without a public record. Users can also turn off Scam Alert at any time if they choose to disable the feature.
WhatsApp's Scam Alert architecture directly contradicts regulators' claims that encrypted messaging platforms cannot offer meaningful safety protections without server-side access to message content. The EU's Chat Control proposal reached its latest flashpoint in July 2026 when EU member states extended a temporary voluntary scanning regime through 2028, but end-to-end encrypted apps were explicitly exempted after Parliament voted 311-228 against encrypted scanning. In India, the Internet and Mobile Association of India (IAMAI) has opposed draft regulations that would require OTT platforms to share data with access providers, arguing that requiring such data sharing amounts to "unconstitutional expropriation of valuable proprietary data." The feature demonstrates that a platform can detect and warn against fraud at scale without a server-side window into message content, challenging the fundamental premise that encryption and platform safety are incompatible. Cloudflare's role as an independent key holder means that even Meta cannot target specific users with different models without creating a public record that researchers can monitor.
The launch addresses significant fraud challenges, with social media fraud losses reaching $2.1 billion in the United States in 2025, representing an eightfold increase from 2020 according to FTC data. WhatsApp ranked second behind Facebook among social platforms by total reported fraud losses, with investment scams accounting for the largest share. The FTC reported that victims lost $425 million to scams via WhatsApp alone out of more than $2.1 billion lost to social media scams in 2025. Globally, the UNODC July 2026 scam report estimated that scam operations across East and Southeast Asia, Australia, and New Zealand cost victims between $88.3 billion and $114.1 billion in 2025 alone. The FBI separately reported that cryptocurrency investment fraud generated $7.2 billion in reported losses from American victims in 2025. WhatsApp removed more than 6.8 million accounts linked to organized scam operations in the first half of 2025, many tied to criminal syndicates operating out of Myanmar, Laos, Cambodia, and the Philippines. However, scammers have adapted quickly to AI-generated lures, making their scripts more linguistically fluent and harder to recognize by ordinary pattern-matching.
WhatsApp is expanding its Bug Bounty programme around the feature, allowing security researchers to examine the system and test whether message content remains on the device. According to SecurityWeek, researchers will receive access to model weights to test whether the model is designed specifically for scam detection and examine how it behaves with different inputs. The company has designed Scam Alert so that Meta or WhatsApp cannot send a particular model to a specific user, with models downloaded from a content delivery network rather than built directly into the app. The feature is off by default and users must enable it manually, with WhatsApp characterizing the current release as an early technical preview rather than a finished product. Meta's bug bounty expansion covering both the model weights and the federated analytics pipeline represents the most concrete accountability mechanism, allowing external researchers to verify that no message content leaves the device and that the model does not exhibit capability beyond scam detection. This launch builds on WhatsApp's broader security initiatives, including warnings for fraudulent device-linking requests introduced earlier this year and the 'Strict Account Settings' feature launched two months earlier to protect journalists and high-risk individuals from sophisticated threats.