BNB $750.45 +1.69%
XRP $1.41 +1.38%
ETH $2,510.39 +1.38%
BTC $82,990.67 +0.83%
BNB $750.45 +1.69%
XRP $1.41 +1.38%
ETH $2,510.39 +1.38%
BTC $82,990.67 +0.83%
BREAKING
Crypto Exchanges

Coinbase Study Reveals New AI Models Fail to Catch Fraud in 16,140 Transactions

Coinbase Finds Newer AI Models Miss More Fraud Across 16,140 Transaction Replay
Coinbase Finds Newer AI Models Miss More Fraud Across 16,140 Transaction Replay

Community Trust ScoreVerified

90%
Real
Verified21 votes
Updated 6 hours ago

Coinbase ran the numbers. And the results were not what the industry expected.

Why It Matters

The findings from Coinbase raise significant concerns about the efficacy of newer AI models in fraud detection, a critical function for maintaining security and trust within crypto markets. As exchanges increasingly rely on advanced technologies to combat sophisticated fraud attempts, this research suggests that upgrading to newer models may not always equate to enhanced performance, potentially exposing platforms to greater risk. This could prompt a reevaluation of AI deployment strategies across the industry, emphasizing the need for rigorous validation processes before adopting newer technologies.

The exchange tested newer versions of three major AI model families against a historical dataset from its Onramp payment service — 16,140 transactions, 7,293 users, 813 confirmed fraudulent payments. Same fixed guidance. Same risk classification policy throughout. The only variable was the model version. And in every single comparison, the newer model caught less fraud than the older one.

Advertisement

Not marginally less. Meaningfully less.

The Benchmark Numbers Are Hard to Ignore

Coinbase compared Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6. Each pairing showed lower recall, a lower F1 score, and lower dollar-weighted recall when the newer version ran. Sonnet took the biggest hit — recall dropped 22.2 percentage points, and dollar-weighted recall fell 22.9 points. That’s not a rounding error. That’s a significant regression in the model’s ability to catch fraudulent money moving through the system.

Opus held up better. Its recall slipped just 0.8 points, which is probably close enough to noise. But GPT’s result was the most interesting and, honestly, the most misleading on the surface. Precision actually went up — 11.5 points higher with GPT-5.6 than GPT-5.4. Sounds like an improvement. But recall dropped 20.7 points, and dollar-weighted recall fell 21.8 points. So the newer model got better at flagging the fraud it did catch, but it flagged far fewer fraud cases overall. More confident. Less complete. That’s a bad trade in payment screening.

The gap between precision and recall matters enormously in fraud detection. A model that’s very precise but misses most fraud is basically useless in a real-money environment. Coinbase’s test made that tension concrete.

The Custom Model That Beat Them All

Here’s where it gets interesting. Coinbase also tested a custom-trained Qwen3.5-9B model, built specifically using historical fraud outcomes. It outperformed Opus 4.5 — which had already outperformed the newer Opus 5 — on several key metrics. F1 score improved by 9.6 points over Opus 4.5. Dollar-weighted recall improved by 35.4 points. And it wasn’t just detection quality. The custom Qwen model cut median end-to-end LLM-request latency by 55% compared to Opus 4.5. Faster and more accurate. That combination matters a lot when you’re screening payments in real time.

The model was trained on Coinbase’s own historical fraud data, with deliberate balancing between fraudulent and legitimate examples. That specificity probably explains most of the performance gap. A general-purpose model trained on broad internet data isn’t optimized for the particular patterns of crypto payment fraud. A specialized model trained on actual fraud outcomes from a specific platform is a fundamentally different tool.

And yet the industry’s default assumption is still pretty much “newer model = better model.” Coinbase’s data pushes back hard on that.

Separately, Coinbase had previously tested adding selective LLM review on top of existing models — not replacing them, just layering on additional review for certain cases. That experiment cut fraudulent transactions by 30% and reduced fraud value by 22%. The recent replay didn’t assess whether newer model versions would produce similar gains, so that comparison can’t be made directly. Gap in the data. Unclear what the newer versions would do in that layered setup.

What Payment Providers Should Take From This

Coinbase’s recommendation is pretty direct: test candidates in your specific decision setup before deploying them. Evaluate latency, reliability, and cost alongside detection quality. Don’t assume a version bump translates to better fraud screening. Run the replay. Check the recall numbers. Check the dollar-weighted recall, not just the headline precision figure.

There’s a real limitation here worth naming. Coinbase’s dataset is proprietary, so SR-Fraud researchers can’t independently replicate the findings. That’s a constraint on how broadly these results generalize. Different platforms, different fraud patterns, different user bases — the numbers could look very different elsewhere. Coinbase can’t really know that, and neither can anyone on the outside.

What the test also didn’t capture: customer losses from actually deploying the newer models. The replay was historical. It can tell you how many of those 813 fraudulent transactions a model would have flagged. It can’t tell you what happens to real users when a regression like Sonnet’s 22.2-point recall drop goes undetected before deployment.

That’s the part that probably keeps the fraud team up at night. The Qwen3.5-9B model’s dollar-weighted recall improvement of 35.4 points over Opus 4.5.

Frequently Asked Questions

What did Coinbase’s AI fraud detection test actually find?

Coinbase found that newer versions of Opus, Sonnet, and GPT models all caught fewer fraudulent payments than their predecessors when tested against 16,140 historical transactions, including 813 confirmed fraud cases, using the same fixed decision policy.

How did the custom Qwen3.5-9B model compare to the other models tested?

The custom Qwen3.5-9B model outperformed Opus 4.5 with a 9.6-point improvement in F1 score and a 35.4-point improvement in dollar-weighted recall, while also cutting median end-to-end LLM-request latency by 55%.

Community Trust IndexHigh Confidence
90%
Real
Real90%10%Fake
21 community signals

Evie Vavasseur

Evie Vavasseur is a crypto writer and digital content specialist covering the latest developments in blockchain technology, decentralized finance, and the broader digital asset ecosystem. With a keen eye for emerging trends, Evie provides accessible and insightful coverage of cryptocurrency markets, NFTs, and Web3 innovations for The Currency Analytics.

Advertisement

Related Stories