DeepSeek-R1 Hallucinates 4x More Than V3, Raising Red Flags for Crypto AI Agent Tokens
DeepSeek-R1, the flagship reasoning mannequin from Chinese lab DeepSeek, hallucinates at 14.3% in accordance with Vectara’s HHEM 2.1 benchmark. That is sort of 4 instances increased than its non-reasoning predecessor DeepSeek-V3, which scored 3.9%. The hole raises arduous questions for the crypto sector. A quick-growing class of AI agent tokens now leans on reasoning-style LLMs…
