Meta has begun testing a new optional Scam Alert feature for WhatsApp that uses on-device machine learning to identify potentially fraudulent conversations. The limited beta rollout comes shortly after a confirmed WhatsApp scam campaign demonstrated that attackers could gain access to accounts without requiring a user’s password.
Meta announced the feature on Aug. 12 through its engineering team, describing it as an additional security layer designed to work alongside WhatsApp’s end-to-end encryption. The company said Scam Alert is being introduced to a limited group of beta users and that it is working with the WhatsApp bug bounty community to test the system under challenging conditions. Meta has not yet announced when the feature will become broadly available or explained how users can join the limited beta.
The system is designed to keep message content on the user’s device. Meta said, “No message content leaves the device for classification or is auto-reported to WhatsApp, Meta, or anyone else.” Once a user chooses to enable Scam Alert, the machine-learning model is downloaded directly to the device, where it evaluates incoming messages from people who are not saved as contacts.
Rather than sending conversations to Meta’s servers for analysis, the model looks for patterns associated with known scams. Meta said the system was trained using scam examples reported by WhatsApp users and uses conversational structure and linguistic signals to make a probabilistic assessment. The company said the feature does not need to automatically share message content with anyone to determine whether a conversation could be fraudulent.
When the system determines that a conversation is likely to involve a scam, it presents a warning to the recipient. The alert is not shown to the person who sent the message, and the recipient can choose whether to block the sender, report the conversation or continue chatting.
WhatsApp also gives users a way to correct an inaccurate warning. If a user believes a conversation was incorrectly identified, they can mark the chat as trusted. Meta said the warning will then disappear and Scam Alert will no longer flag that particular conversation again.
The new feature adds another layer to WhatsApp’s broader defense-in-depth security strategy, sometimes described as the Swiss Cheese security model. The approach relies on multiple independent protections so that a weakness in one layer does not necessarily allow an attacker to reach the user.
WhatsApp already offers other safeguards aimed at preventing account and device-based attacks. A recent security tools update can warn users about potential attacks involving device-linking codes, while the optional Strict Account Settings configuration can automatically silence calls from unknown numbers, block attachments from unknown users and prevent link previews.
Meta’s approach follows similar efforts elsewhere in the technology industry. Google has introduced AI-based protections designed to detect scams involving calls and messages, including a fake-call detection feature for users of Phone by Google when both participants are using the service. Google’s Android protections also perform AI analysis directly on the device while keeping conversations private.
Leave a comment