@axelis
Gave Kimi K2 Thinking Quest 3 of #randombench
and it is completely stuck in a loop and couldn't unstuck itself.
Ernie 5.0 Preview also got stuck and crashed out all via @arena
Quest 3 is a true test of these models and only 2 models reliable PASS it in one shot
1. Grok 4
2.GPT-5 High
Other variants of the two models like any thinking version of gpt-5 and grok 4 fast and even 03, 04 only scores it if the test is run in their chat UI which means a sort of helpers like code interpreter or python scripts could be aiding them. Pure API would fatally render them unable to PASS .
The same is happening to Kimi K2 Thinking, Qwen 3 Max thinking and Ernie 5.0 preview, they cannot reliably one-shot and pass the challenge via pure API.
we can conclude that only 2 models can solve the problem. Even gpt-5 High and Grok 4 can fail sometimes! But they have a better PASS rate than any other models.
https://x.com/Web3Aible/status/1986879725667488181?t=QHwFJLoIMLw_vqYweYIFUQ&s=19