EvilTokens phishing-as-a-service platform diagram showing compromised Microsoft 365 inboxes across organizations

Microsoft obtained a US federal court order to seize 50 websites and more than 150 associated domains tied to EvilTokens, a phishing-as-a-service platform that had been operating since February 2026. The seizure was part of a coordinated takedown with industry and law enforcement partners. UK police arrested two men, ages 32 and 38, suspected of running the technology and infrastructure behind the service, and both have been released on police bail pending forensic examination of seized devices. Details in this article come from analysis published by CSO Online.

EvilTokens was a subscription product, not a one-off campaign. Criminals paid a $1,500 sign-up fee and $500 per month for a dashboard and chatbot that bundled account compromise, mailbox analysis, target selection, and fraud preparation into one interface, marketed through Telegram channels.

Microsoft linked the platform to more than 12,000 compromised Microsoft 365 inboxes across more than 10,000 organizations worldwide.

The technical core was device code phishing, which abuses Microsoft's own OAuth 2.0 device-code authentication flow. A victim clicked a link, saw a short-lived authentication code, and entered it on the genuine Microsoft device login page. That handed the operator a valid session token without the victim's password ever being exposed, which is why this style of compromise reaches accounts that already have multi-factor authentication in place.

Where this stops being a niche identity exploit is the AI layer. EvilTokens included an "analyst" chatbot that read compromised mailboxes to find who controls payments, which business relationships carry trust, and which invoices were in flight, then recommended impersonation targets and helped draft fraudulent messages grounded in real conversations. That is the machinery of business email compromise, packaged for buyers with no prior skill in identity attacks or financial fraud.

Victims spanned wholesale distribution, construction, financial services, real estate, higher education, and healthcare across North America, the UK, France, India, and Australia. Coinbase traced roughly $1.1 million in revenue across more than 700 distinct crypto addresses connected to the operation. Microsoft has not publicly attributed EvilTokens to a named criminal group beyond the two arrests.

How EvilTokens Abused AI Services and Stolen Credentials

The technical core of EvilTokens was the abuse of Microsoft's OAuth 2.0 device-code authentication flow, a legitimate mechanism built for devices that cannot display a browser login page, such as smart TVs and conference room hardware. A user is shown a short code and told to enter it at a genuine Microsoft device login page. EvilTokens generated those codes on the attacker's side and delivered them through tailored phishing messages, so the victim authenticated on real Microsoft infrastructure while the resulting session token was issued to the criminal.

That distinction matters operationally. The password is never typed into a fake page, never captured, and never changed. MFA completes normally because the user genuinely approves their own sign-in. Maps to MITRE ATT&CK T1566 (Phishing) for delivery and T1528 (Steal Application Access Token) for the outcome, with T1550.001 (Use Alternate Authentication Material) covering how the stolen token was replayed against Microsoft 365 and Entra ID.

Persistence came from token refresh. Rather than re-phishing the victim, the platform continued renewing access tokens and monitoring the mailbox, which kept the criminal inside the account after the original short-lived code had expired. For your business, that means a password reset alone does not evict the intruder, and the compromise can sit inside a mailbox that shows no failed logins and no password change events.

The AI layer sat on top of that access. Microsoft describes an "analyst" chatbot that read compromised inboxes and produced fraud opportunities from what it found. Jason Rivera, global field CISO at SimSpace, described the analysis as identifying who controls payments, which business relationships carry trust, and which invoices or transactions present opportunities.

Rivera outlined the downstream steps the platform automated:

  • Recommending impersonation targets based on real correspondence
  • Drafting fraudulent messages grounded in actual business conversations
  • Running automated reconnaissance to map organizational permissions
  • Maintaining access through token refresh and continuous inbox monitoring

In ATT&CK terms that covers T1114 (Email Collection), T1087 (Account Discovery), and T1534 (Internal Spearphishing). The business consequence is speed. A criminal with no accounting background could read a mailbox and produce a credible payment-diversion message in the language and formatting the recipient already expects, which is what makes business email compromise succeed against finance teams who do check for odd phrasing.

Omair Manzoor of ioSENTRIX framed the timing shift bluntly, advising that organizations assume every compromised mailbox will be read and exploited by AI within minutes rather than days. That compresses the window between initial token theft and an attempted fraudulent wire.

The reporting leaves several details unaddressed. Microsoft's account does not name the underlying AI model or provider behind the analyst chatbot, does not describe any jailbreak or guardrail-bypass technique used to obtain that capability, and does not indicate whether stolen API keys funded the inference. There is no published evidence in this reporting of a reverse-proxy or model-endpoint abuse layer of the kind seen in other AI-assisted crime services. Beyond the more than 12,000 compromised inboxes and the sectors hit, which ranged from wholesale distribution and construction through financial services, real estate, higher education, and healthcare, the source provides no file hashes, sender patterns, or infrastructure fingerprints. Treat the device-code flow itself as the reliable technical indicator here.

What AI-Powered Cybercrime Means for Organisations

Microsoft counted more than 12,000 compromised Microsoft 365 inboxes in this operation, spread across wholesale distribution, construction, financial services, real estate, higher education, and healthcare, with victims in North America, the UK, France, India, and Australia. The spread across unrelated sectors matters for you because the selection criterion was not industry. It was a mailbox containing payment conversations.

Omair Manzoor of ioSENTRIX advised organisations to assume every compromised mailbox will be read and exploited by AI within minutes rather than days. That compresses the window between a user approving a login code and a fraudulent invoice being drafted from your own correspondence.

A stolen session token carries every entitlement the account holds. Where your Microsoft 365 identities also unlock paid AI services, assistant seats, or metered API access, an attacker holding that token consumes those services under your billing. The charges surface on your invoice as ordinary usage by a real employee, which makes them slow to spot and awkward to dispute with the provider afterwards.

The reputational exposure is separate from the money. Jason Rivera of SimSpace described the platform recommending impersonation targets and drafting fraudulent messages grounded in actual business conversations. Your customers and suppliers receive mail from a genuine colleague, referencing a genuine invoice thread, sent through genuine Microsoft infrastructure. When the fraud is uncovered, the counterparty's loss is attached to your domain and your staff member's name.

Microsoft's own description of the platform is the part worth taking to your board: capabilities that previously required experience across identity attacks, cloud systems, social engineering, and financial fraud were packaged into a ready-made interface. The practical consequence for you is that the quality of a phishing message no longer indicates the skill of the sender. Grammar errors, generic greetings, and mismatched branding are no longer reliable signals for your staff, and the volume of credible attempts rises alongside the quality.

Governance consequences follow quickly. Acceptable-use terms on AI and cloud services generally hold the account owner responsible for what the account produces. If abusive or criminal content is generated using your organisation's credentials, the provider may suspend the account or restrict the tenant while it investigates, which affects staff who had no involvement in the incident. You then explain that suspension to customers who cannot reach you by email.

Data exposure runs deeper than the fraud attempt. Once attacker tooling reads a mailbox, contract negotiations, payroll threads, patient or student correspondence, and board material all become inputs to someone else's analysis. For your UK and French operations that engages GDPR notification duties, and healthcare and higher education carry their own sector obligations. Those duties attach to the exposure itself, whether or not any payment was redirected.

API keys and tokens held by individual staff or by third parties are the accountability gap your auditors will probe. When a contractor's key or a developer's personal token is used to generate fraudulent content, the billing record and the audit trail still resolve to your tenant. Where ownership of those credentials was never documented, establishing who approved the access, what it could reach, and whether it is still live becomes an investigation in its own right.

The direct loss in these cases is the redirected payment. The larger cost is the reconstruction work afterwards: determining which mailboxes were read, which conversations were copied, which counterparties received messages under your name, and which of them need to be told.

Detecting Token Abuse and Securing AI Service Access

Start with your Entra ID Conditional Access policy and block the device code flow for every user who does not need it. A grant restricting device code authentication removes the mechanism the operators depended on, and in most organisations the legitimate users of that flow are a short list of shared conference room devices and a handful of IoT endpoints. Scope an exception group for those, and deny the flow everywhere else.

Then pull an inventory of the AI service credentials your teams already hold. That means API keys for assistants and copilots, service principals with mailbox or Graph permissions, and OAuth application consents users granted without review. Revoke anything unused or long lived, because a key issued a year ago with no expiry gives an intruder access that outlasts any password reset you perform.

  • Review billing and usage dashboards on every AI subscription for spikes, off-hours consumption, or calls originating from regions where you have no staff.
  • Enforce MFA on the administrative accounts that own those subscriptions, including the billing owner, which is frequently a finance or procurement login rather than an IT one.
  • Revoke active refresh tokens for any account you suspect, since disabling a password leaves an issued session valid until it expires on its own.
  • Shorten token lifetimes in your sign-in frequency policy so a stolen session ends in hours instead of persisting for weeks.

For detection, filter your sign-in logs on the device code authentication protocol and treat every result as something to explain. Pair that with alerting on token refreshes from a new autonomous system number, new inbox rules that move or delete messages containing invoice keywords, and consent grants to applications nobody in IT registered. In environments Capstone manages, Adlumin correlates these authentication anomalies against normal user behaviour, flagging a session that suddenly refreshes from an unfamiliar network while the same account is active elsewhere.

Over the next quarter, put a rotation schedule on every AI service key and scope each one to least privilege. A key that only needs to read a document set should not carry write access to a mailbox. Turn on logging at the AI endpoint itself so you have prompt and API call records to investigate against, since without them you cannot establish what an intruder asked the assistant to summarise.

Scan your code repositories and CI/CD pipelines for committed keys. Developers embed them in build configuration and test scripts, and a public or loosely permissioned repo turns one convenience shortcut into standing access to your AI tenant. Automated secret scanning on push, plus a one-time historical scan of commit history, catches the bulk of it.

The longer term work is governance. Write an AI acceptable-use policy that states which services staff may connect to company data and which require approval, then centralise procurement so finance can see every AI subscription being paid for. Shadow subscriptions bought on a departmental card are the ones nobody rotates keys for.

Finally, subscribe to credential exposure monitoring covering your domains, including API key formats as well as email addresses. The operators behind services like this one sell access before they use it, and finding your own credential listed gives you the chance to revoke it before a buyer acts. Assign one person ownership of that feed so alerts get triaged rather than filed.

Treating AI Credentials as Critical Assets

Takedowns of this kind disrupt an operation without ending the business model behind it. The two men arrested in the UK were released on police bail while their seized devices are examined, and the subscriber base that paid for the service has not been identified or charged. Omair Manzoor of ioSENTRIX attributed the disruption to a specific operational error by the operators, namely centralizable domains and traceable crypto payments, and he expects more sophisticated versions of the same scam to follow.

That matters for how you plan. The next service to occupy this niche will likely distribute its infrastructure more widely and handle payments in ways that resist tracing, which means you should treat court-ordered domain seizures as relief rather than as a change in your risk profile. Your exposure comes from the authentication design that made the attack work, and that design is still in place.

The practical takeaway is a classification decision. Treat every API key and authentication token that grants access to an AI service the way you already treat domain administrator credentials: a named owner, a recorded expiry, a documented business justification, and an entry in your asset register. Most organisations track these in a spreadsheet, if at all, because they were issued by developers or business units experimenting with assistants and copilots.

Watch their usage with the same attention you give privileged account activity. A key that reads mailboxes on behalf of an AI service has effectively the same reach as the session tokens this operation stole, and it should carry the same level of oversight in your environment.

In This Article

Top hits