Newer AI models missed more payment fraud in Coinbase’s benchmark
Coinbase reported Oct. 7 that newer variations of three main AI mannequin households caught fewer fraudulent funds and a smaller share of fraud worth in a historic take a look at of payment screening for its Onramp service, regardless of an unchanged resolution coverage. The findings problem the idea that upgrading a mannequin improves an current payment screener.
The firm’s evaluation replayed 16,140 transactions throughout 7,293 customers, together with 813 confirmed fraudulent transactions. The cohort lined 9 weeks earlier than its danger agent rolled out, retaining all matured fraud circumstances whereas sampling respectable site visitors.
Each candidate reviewed latest transaction habits underneath mounted steering and the identical coverage for turning danger classifications into choices. This remoted the choice mannequin’s habits inside that setup, fairly than evaluating redesigned screening methods.
Results from a hard and fast historic replay
Coinbase in contrast Opus 4.5 with Opus 5, Sonnet 4.6 with Sonnet 5, and GPT-5.4 with GPT-5.6 (sol). Every newer model had decrease recall, a decrease mixed precision-and-recall rating referred to as F1, and decrease dollar-weighted recall. Recall measures the share of fraud circumstances a mannequin catches; dollar-weighted recall measures how a lot of the full fraud worth it catches.
Sonnet’s recall fell 22.2 proportion factors and its dollar-weighted recall dropped 22.9 factors. Opus’s recall declined 0.8 factors. Both newer models additionally had decrease precision, that means a smaller share of transactions they categorized as fraud had been really fraudulent.
GPT confirmed why one bettering rating could be deceptive. Its precision rose 11.5 proportion factors, however recall fell 20.7 factors and dollar-weighted recall fell 21.8 factors. Its fraud flags had been more correct, whereas more fraud circumstances and worth escaped detection in the replay.
The replay doesn’t set up buyer losses from deploying these variations. Coinbase additionally stated it might establish the regressions with out establishing their trigger.
Coinbase’s earlier online experiment in contrast including selective LLM evaluate with the present models and guidelines alone. That agent-enabled stream recorded 30% fewer fraudulent transactions and 22% much less fraud worth; it didn’t examine newer mannequin variations.
In their limitations, the SR-Fraud researchers say the proprietary dataset can’t be launched, limiting unbiased replication and generalization. Their related payment-fraud study first appeared Sept. 23 and was revised Sept. 30, earlier than the October blogs.
A separate case for a customized mannequin
In its Oct. 8 disclosure, Coinbase reported {that a} post-trained Qwen3.5-9B mannequin exceeded Opus 4.5 throughout 4 fraud-detection metrics. F1 improved 9.6 proportion factors and dollar-weighted recall rose 35.4 factors. The firm specialised it utilizing historic fraud outcomes and deterministic rewards balancing fraudulent and bonafide examples.
Separately, manufacturing measurements put median end-to-end LLM-request latency at 0.683 seconds versus 1.515 seconds for Opus 4.5, a 55% relative discount. Faster inference and stronger benchmark detection got here from totally different evaluations.
For payment suppliers, the improve query is whether or not a candidate improves fraud protection underneath their precise resolution setup. Coinbase recommends testing that configuration first, then evaluating modified prompts or thresholds individually, with latency, reliability and value alongside detection high quality.
The put up Newer AI models missed more payment fraud in Coinbase’s benchmark appeared first on CryptoSlate.

