WhatsApp Introduces Scam Alert Feature to Combat Social Engineering Attacks

WhatsApp has unveiled a new optional feature called Scam Alert, designed to protect users from potential scam messages by leveraging on-device machine learning. This initiative aims to address the growing sophistication of social engineering attacks, including those enhanced by AI-generated content, while maintaining the platform’s commitment to end-to-end encryption.

When enabled, Scam Alert downloads a lightweight machine learning model directly to the user’s device. This model analyzes incoming messages from unknown contacts, identifying conversational structures and linguistic patterns commonly associated with scams. Importantly, all processing occurs locally on the device, ensuring that message content remains private and is not transmitted to WhatsApp, Meta, or any third parties unless the user chooses to report a message.

If a message is flagged as a potential scam, the recipient receives a discreet warning within the chat interface. The user can then decide to block the sender, report the message, continue the conversation, or mark the chat as trusted if they believe the warning is a false positive. This approach emphasizes user autonomy and privacy, aligning with WhatsApp’s principles of on-device processing, no automatic reporting, and full user control.

To assess the effectiveness of Scam Alert without compromising user privacy, WhatsApp has developed a confidential federated analytics system. This system utilizes Trusted Execution Environments (TEEs) to aggregate anonymous data, such as the number of warnings displayed and user actions taken, while applying differential privacy techniques to ensure individual data points remain untraceable. This method allows WhatsApp to gather performance metrics without accessing personal message content.

Addressing potential security concerns, WhatsApp has implemented measures to prevent targeted delivery of manipulated models. Each version of the Scam Alert model is published with a SHA-256 hash to a third-party append-only transparency ledger before deployment. Additionally, download requests are anonymized through an Oblivious HTTP (OHTTP) relay, and experiment group assignments for testing new model variants are conducted entirely on-device using locally generated randomness. These steps ensure that neither external attackers nor malicious insiders can manipulate the model delivery process.

Users have the ability to audit the system through an in-app transparency log, accessible via the Account settings under ‘Request Info’ and ‘Scam Alert Activity.’ This log provides information on which messages were scanned and the model version used, promoting transparency and trust in the feature’s operation.

WhatsApp is also expanding its Bug Bounty program to include the Scam Alert model and the federated analytics pipeline, inviting external researchers to verify the system’s integrity and effectiveness. This collaborative approach aims to identify and address potential vulnerabilities proactively.

Scam Alert is initially being rolled out in a limited beta phase, allowing WhatsApp to refine the system based on user feedback and further testing. The company plans to publish a detailed engineering white paper on the pipeline’s design, building upon its previous work presented at the USENIX NSDI 2025 conference.

As social engineering tactics become increasingly sophisticated, features like Scam Alert represent a crucial step in enhancing user security. By combining advanced machine learning with robust privacy safeguards, WhatsApp aims to provide users with effective tools to identify and mitigate potential scams without compromising their privacy. This development underscores the importance of continuous innovation in cybersecurity to stay ahead of evolving threats.