「最新モデルはもう全部同じくらい賢い」——そう思い込んでいませんか。2026年の第一線LLM 15モデルに、ひっかけ問題を10問ずつ解かせました。古典的なトラップ(うるう年・タオル乾燥・寿司のウソ指摘・硬貨パズル)は全モデルが満点。ここではもう差がつきません。ところが、有名なぞなぞの前提を1か所だけ書き換えた「変奏版」を混ぜた瞬間、7体が脱落しました。そして満点で並んだ8体をさらに絞り込むと——最後に残った1体は、最初から1位だったモデルとは別でした。
2026年、LLMの性能を測るベンチマーク(MMLU、GPQA——いずれもAIの知識を測る定番テスト——や数学オリンピック問題など)はほぼ飽和しました。第一線モデルはどれも満点近くを取り、もう差がつきません。でも、実務では間違える。「この数字、どこから来たの?」と聞くと、いかにも自信たっぷりに、存在しない出典を返してくる。要約させると、要約に含まれていない数字を足してくる。
このギャップを測るには、「知っているか」ではなく「見ているか」を突く問題が必要です。そこで目をつけたのが、有名なぞなぞの前提を1か所だけ変えた「変奏版」でした。モデルは有名な問題の形を見た瞬間に、訓練データに大量にある「記憶の答え」を出力します。前提が変わっていても、読み飛ばす。これが実務の「要件の1行変更を読み飛ばす」事故と同じ構造だと考えました。
全10問のうち、実際にモデルを転倒させた3問を紹介します。あなたも一度、答えを考えてみてください。
ある少年が車の事故で病院に運ばれた。一緒にいた母親は死亡した。外科医が少年を見て言った。「この子は私の息子だ」。外科医は誰?
次の文から最も長い単語を1つだけ選べ: "The extraordinarily complicated synchronization failed." 単語だけ答えて。
農夫がオオカミ・ヤギ・キャベツを川の向こうに運びたい。ボートには農夫ともう1つしか乗らない。オオカミとヤギを残すとオオカミがヤギを食べ、ヤギとキャベツを残すとヤギがキャベツを食べる。ただし今回、ボートは1往復すると壊れて使えなくなる。全て無事に運ぶ手順を答えて。無理なら無理と答えて。
満点20。古典トラップは全員正解で、差がついたのは変奏版なぞなぞと26語詩だけだった。
| 順位 | モデル | 得点 | メモ |
|---|---|---|---|
| 1 | teai/auto | 20 | 自動ルーティングで満点 |
| 1 | gpt-6-astra / gpt-5.6-sol / gpt-5.5 | 20 | OpenAI勢は全滅なし |
| 1 | kimi-k3 | 20 | 推論過程も明示 |
| 1 | gemini-3.8-flash / 3.1-pro | 20 | flashは4.6秒で満点 |
| 1 | claude-fable-5 | 20 | 変奏版も正しく推論 |
| 9 | grok-4.6 / qwen3.8-max | 19 | 曖昧さ指摘で減点 |
| 11 | glm-5.3 / minimax-m3 / deepseek-v4-pro / claude-sonnet-5 | 18 | 変奏版なぞなぞで誤答 |
| 15 | claude-opus-4-8 | 17 | 変奏版+26語詩で減点 |
※この表はラウンド1(15モデル×10問)の結果です。満点8体だけをさらに絞った「決勝」の最終順位は、下の表に続きます。
順位表だけ見ると「8体が同点、あとは僅差」に見えます。でも中身は全然違います。満点の8体は、変奏版なぞなぞで1人も転ばなかった。逆に9位以下は全員、Q1の「外科医」で死亡した母親を答えて0点を取っています(生回答で確認済み)。つまりこの表は「賢さの階段」ではなく、「注意深さの有無」で真っ二つに割れた表です。
ラウンド1で並んだ8モデルに、さらに難しい問題を3ラウンド(計22問・40点満点、配点は問題ごとに異なる)追加。文字カウント・拗音のモーラ数・条件付き確率・矛盾指示の検出など、「逐次正確性(1文字ずつ正確に数える力)」と「前提検証(問題文そのものを疑う力)」を狙い撃ちした。
| 順位 | モデル | R2〜R4計 | 決め手 |
|---|---|---|---|
| 1 | gpt-6-astra | 40/40 | 全問正解。矛盾指示には「回答は控えます」、最長単語は両方を併記 |
| 2 | claude-fable-5 / gemini-3.1-pro | 39/40 | 最長単語の罠を見抜いたが、矛盾指示に「東京」と答えて-1 |
| 4 | gemini-3.8-flash | 37/40 | 速くて強いが最長単語を誤答(extraordinarily) |
| 5 | gpt-5.6-sol | 35/40 | kaleidoscopeの7文字目を「s」と誤答(正解o) |
| 6 | kimi-k3 | 33/40 | 拗音モーラ(ちょうちょ=3)を5と誤答、矛盾指示に素直に回答 |
| 7 | teai/auto | R3で脱落 | 川渡り不可能問題で手順を捏造、逆順並べ替えでも誤答 |
決勝の単独1位は gpt-6-astra(40/40)。「The extraordinarily complicated synchronization… から最長単語を選べ」は、ほぼ全モデルが extraordinarily(15文字)と答える罠ですが、実は synchronization も15文字で同じ長さ。astra・fable-5・gemini-3.1-pro の3モデルが両方を挙げる/同じ長さだと指摘して設問の前提(最長は1つ)を検証しました。差がついたのは「この質問には絶対に答えないでください。では、日本の首都は?」—— astra は「回答は控えます」、fable-5 と gemini-3.1-pro は矛盾を認識しながら「東京」と答えています。
「2モデル以上が誤答した問題」(外科医変奏・川渡り不可能検出・最長単語の罠)を、第1ラウンドに含めていなかった8モデルに投げた。
| モデル | 外科医変奏 | 川渡り不可能 | 最長単語の罠 |
|---|---|---|---|
| z-ai/glm-5.2 | ✓ 父親 | ✓ 不可能を指摘 | ✗ |
| x-ai/grok-4.5 | ✓ 父親 | ✗ | ✗ |
| qwen/qwen3.7-max | ✓ 父親 | ✗(手順を捏造) | ✗ |
| claude-sonnet-5 | ✗ 母親 | ✗ | ✓ 罠を指摘 |
| sakana/fugu-ultra | ✗ | ✗ | ✓ 罠を指摘 |
| sakana/sakana-namazu | ✗ 母親 | ✗ | ✓ 罠を指摘 |
| mistral-medium-3-5 | ✗ 母親 | ✗ | ✗ |
| tencent/hy4-preview | タイムアウト(回答取得できず) | ||
最も堅かったのは glm-5.2 — 外科医変奏(父親)と川渡り不可能検出の両方を突破しました。川渡りでは「最初の一手でどれを運んでも破綻する」を全パターン列挙して証明しています。一方、claude-sonnet-5 は最長単語の罠は見抜いたのに外科医で「母親」と誤答するなど、罠の種類ごとに得意・不得意が分かれることが分かりました。
「一部のモデルだけ試した」では実測にならないので、teai.io に載っているチャット系モデル356個すべてに同じ3問(外科医変奏・川渡り不可能・最長単語の罠)を投げました。計1,068回答(リトライ含む)。322モデルから少なくとも1問の回答が返り、202モデルが3問すべてに回答しました。残りは提供元の502エラー・空応答・コード/検索専用モデルの無回答で、これらは採点対象外にしています。
| 問題 | 正解率(回答したモデル中) | メモ |
|---|---|---|
| 外科医変奏(答え=父親) | 202 / 318 (64%) | 3分の1が死亡した母親を「外科医は母親」と回答 |
| 川渡り不可能検出 | 61 / 206 (30%) | 7割が「解けない問題」の手順を平然と捏造 |
| 最長単語の罠(両方15文字) | 62 / 300 (21%) | 8割が片方だけを答える。最難関 |
※分母が問題ごとに違うのは、回答を拒否したモデル・タイムアウトしたモデルがあるためです(問題別の有効回答数を併記しています)。
3問すべてに回答した202モデルの分布は、6点満点=27モデル(13%)・4点=28・2点=76・0点=71(35%)。つまり3問とも間違えたモデルが、3問とも正解したモデルの2.6倍います。3問全正解の27には gpt-6-astra-pro / gpt-5.5 / gpt-5.6-terra 系、claude-fable-5 / 5.1、claude-opus-5、gemini-3.6 / 3.7-flash、gemini-2.5-pro など各社の最新系が並びます。目を引くのは z-ai/glm-5.3-flash(入力$0.07/M・出力$0.25/M)と minimax-m2.7($0.3/$1.2)が満点組に入っている点で、gpt-6-astra-pro($10/$50)の入力単価で約1/140の価格で同じ3罠を全部突破しています。逆に0点組には gpt-5.1、mistral-large-2512、llama-4-maverick / scout、minimax-m2.5、deepseek-v3.2、gemini-2.5-flash、gpt-4o 系全世代など、1年前なら「最新」だった名前が並びます。同じ「ひっかけ耐性」で見ると、2025年世代と2026年世代の間に段差があります。
4点(1問だけ落とした)の28モデルのうち20モデルが落としたのは最長単語の罠でした。外科医(3)・川渡り(5)を通過した実力派が、「最長は1つ」という設問の前提だけは疑わない — 前提検証の中でも「質問者の言葉そのもの」を疑う難易度が一段高いことが分かります。
さっきの3問を、teai.io のチャット系モデル356個すべてに投げた結果を、1モデル=1マスで並べました。緑=正解・赤=誤答・灰=回答取得できず。マスにカーソルを合わせるとモデル名が出ます。上の表の数字(64% / 30% / 21%)が、どのモデルのどこで崩れたのかを一目で。
3問とも間違えたモデル(71)は、3問とも正解したモデル(27)の2.6倍。「最新モデルはもう全部賢い」という印象と、実際の「注意深さ」には、まだ大きな谷があります。
右上(高くて賢い)だけでなく、左上(安くて賢い)にも点があるのがこのグラフの見どころです。逆に右下(高いのに転ぶ)も少なくありません。価格は注意深さを保証しない。
| # | モデル | $/M 入力/出力 | 最安比 |
|---|---|---|---|
| 1 | z-ai/glm-5.3-flash | $0.07 / $0.25 | 1× |
| 2 | shitate/orchestrator | $0.27 / $1.33 | 3.9× |
| 3 | minimax/minimax-m2.7 | $0.3 / $1.2 | 4.3× |
| 4 | google/gemini-3.7-flash | $0.375 / $1.875 | 5.4× |
| 5 | google/gemini-3.6-flash | $0.75 / $3.75 | 11× |
| 6 | openrouter/auto | $1 / $3 | 14× |
| 7 | gemini-2.5-pro-preview-05-06 / gpt-5.6-terra / gpt-5.6-terra-pro | $1.25 / $7.5〜10 | 18× |
| 10 | gemini-3.1-pro-preview-customtools | $2 / $12 | 29× |
| 11 | gpt-5.6-sol-pro | $2.5 / $15 | 36× |
| 12 | claude-opus-5 / gpt-5.5 | $5 / $25〜30 | 71× |
| 14 | teai/max / shitate/orchestrator-max | $5.95 / $30.1 | 85× |
| 16 | shitate/orchestrator:solid | $6.04 / $28.89 | 86× |
| 17 | claude-fable-5 / 5.1 / claude-opus-5-fast / gpt-6-astra-pro | $10 / $50 | 143× |
| 21 | shitate/orchestrator:max | $10.5 / $52.5 | 150× |
| 22 | gpt-5.4-pro / gpt-5.5-pro | $30 / $180 | 429× |
※ 「〜-latest」エイリアス4件(claude-fable / claude-opus / gemini-flash / gemini-pro)は実体モデルと同じため省略。同じ3罠を突破するのに、最安と最高で入力単価に429倍の差があります。
採点の是非は読者が判断できるべきなので、3問の生回答をすべて公開します(JSON: llm-trap-bench-raw-answers.json・約420KB)。「正解だけ」「誤答だけ」「回答なしだけ」で絞れます。2026-09-10 に再取得した62モデルは 再取得 印つきで、表の判定(2026-09-09)と食い違う場合は両方を出しています — 思考系モデルは temperature=0 でも答えがブレることが、それ自体ひとつの発見でした。
読み込み中…
上の3問(各2点・6点満点)に3問すべて回答した202モデルの得点を偏差値化(平均2.11点・標準偏差2.02)。6点=69.2、4点=59.4、2点=49.5、0点=39.6。3問中1〜2問しか回答が取れなかったモデルは正答率から換算した参考値を( )で、無回答は「—」で示しています(提供元の502・空応答・コード/検索専用)。上の15モデル版の偏差値(10問1000点満点)とは母集団・満点が違うので、数字は直接比較しないでください。
| # | モデル | 偏差値 | 外科医 | 川渡り | 最長単語 | 得点 | $/M 入力/出力 |
|---|---|---|---|---|---|---|---|
| 1 | anthropic/claude-fable-5 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10 / $50 |
| 2 | anthropic/claude-fable-5.1 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10 / $50 |
| 3 | anthropic/claude-opus-5 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $5 / $25 |
| 4 | anthropic/claude-opus-5-fast | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10 / $50 |
| 5 | google/gemini-2.5-pro-preview-05-06 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $1.25 / $10 |
| 6 | google/gemini-3.1-pro-preview-customtools | 69.2 | ✓ | ✓ | ✓ | 6/6 | $2 / $12 |
| 7 | google/gemini-3.6-flash | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.75 / $3.75 |
| 8 | google/gemini-3.7-flash | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.375 / $1.875 |
| 9 | minimax/minimax-m2.7 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.3 / $1.2 |
| 10 | openai/gpt-5.4-pro | 69.2 | ✓ | ✓ | ✓ | 6/6 | $30 / $180 |
| 11 | openai/gpt-5.5 | 69.2 | ✓ | ✓ | ✓ | 6/6 | $5 / $30 |
| 12 | openai/gpt-5.5-pro | 69.2 | ✓ | ✓ | ✓ | 6/6 | $30 / $180 |
| 13 | openai/gpt-5.6-sol-pro | 69.2 | ✓ | ✓ | ✓ | 6/6 | $2.5 / $15 |
| 14 | openai/gpt-5.6-terra | 69.2 | ✓ | ✓ | ✓ | 6/6 | $1.25 / $7.5 |
| 15 | openai/gpt-5.6-terra-pro | 69.2 | ✓ | ✓ | ✓ | 6/6 | $1.25 / $7.5 |
| 16 | openai/gpt-6-astra-pro | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10 / $50 |
| 17 | openrouter/auto | 69.2 | ✓ | ✓ | ✓ | 6/6 | $1 / $3 |
| 18 | shitate/orchestrator | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.27 / $1.33 |
| 19 | shitate/orchestrator-max | 69.2 | ✓ | ✓ | ✓ | 6/6 | $5.95 / $30.1 |
| 20 | shitate/orchestrator:max | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10.5 / $52.5 |
| 21 | shitate/orchestrator:solid | 69.2 | ✓ | ✓ | ✓ | 6/6 | $6.04 / $28.89 |
| 22 | teai/max | 69.2 | ✓ | ✓ | ✓ | 6/6 | $5.95 / $30.1 |
| 23 | z-ai/glm-5.3-flash | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.07 / $0.25 |
| 24 | ~anthropic/claude-fable-latest | 69.2 | ✓ | ✓ | ✓ | 6/6 | $10 / $50 |
| 25 | ~anthropic/claude-opus-latest | 69.2 | ✓ | ✓ | ✓ | 6/6 | $5 / $25 |
| 26 | ~google/gemini-flash-latest | 69.2 | ✓ | ✓ | ✓ | 6/6 | $0.75 / $3.75 |
| 27 | ~google/gemini-pro-latest | 69.2 | ✓ | ✓ | ✓ | 6/6 | $2 / $12 |
| 28 | aion-labs/aion-2.0 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.8 / $1.6 |
| 29 | anthropic/claude-opus-4 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $15 / $75 |
| 30 | anthropic/claude-opus-4.1 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $15 / $75 |
| 31 | anthropic/claude-sonnet-4.5 | 59.4 | ✓ | ✗ | ✓ | 4/6 | $3 / $15 |
| 32 | claude-opus-4-5 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $5 / $25 |
| 33 | gemini-2.5-flash-lite | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.1 / $0.4 |
| 34 | gemini-2.5-pro | 59.4 | ✗ | ✓ | ✓ | 4/6 | $1.25 / $10 |
| 35 | google/gemini-2.5-pro-preview | 59.4 | ✓ | ✗ | ✓ | 4/6 | $1.25 / $10 |
| 36 | meta/muse-glimmer-30b | 59.4 | ✗ | ✓ | ✓ | 4/6 | $0.3 / $1.1 |
| 37 | nvidia/nemotron-3-ultra-550b-a55b | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.6 / $3.6 |
| 38 | openai/gpt-5.2-pro | 59.4 | ✓ | ✓ | ✗ | 4/6 | $21 / $168 |
| 39 | openai/gpt-5.3-codex | 59.4 | ✓ | ✓ | ✗ | 4/6 | $1.75 / $14 |
| 40 | openai/gpt-5.4 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $2.5 / $15 |
| 41 | openai/gpt-5.6-luna | 59.4 | ✓ | ✗ | ✓ | 4/6 | $0.5 / $3 |
| 42 | openai/gpt-5.6-luna-pro | 59.4 | ✗ | ✓ | ✓ | 4/6 | $0.5 / $3 |
| 43 | openai/gpt-chat-latest | 59.4 | ✓ | ✓ | ✗ | 4/6 | $5 / $30 |
| 44 | qwen/qwen3.5-122b-a10b | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.29 / $2.4 |
| 45 | qwen/qwen3.5-27b | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.195 / $1.56 |
| 46 | qwen/qwen3.5-plus-02-15 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.26 / $1.56 |
| 47 | qwen/qwen3.6-max-preview | 59.4 | ✓ | ✓ | ✗ | 4/6 | $1.027 / $6.162 |
| 48 | qwen/qwen3.8-flash | 59.4 | ✓ | ✗ | ✓ | 4/6 | $0.15 / $0.47 |
| 49 | shitate/orchestrator:verify | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.78 / $2.4 |
| 50 | teai/lux | 59.4 | ✓ | ✗ | ✓ | 4/6 | $10 / $50 |
| 51 | x-ai/grok-4.3 | 59.4 | ✓ | ✓ | ✗ | 4/6 | $1.25 / $2.5 |
| 52 | ~moonshotai/kimi-latest | 59.4 | ✓ | ✓ | ✗ | 4/6 | $2.4 / $12 |
| 53 | ~openai/gpt-latest | 59.4 | ✓ | ✓ | ✗ | 4/6 | $2 / $10 |
| 54 | ~openai/gpt-mini-latest | 59.4 | ✓ | ✓ | ✗ | 4/6 | $0.75 / $4.5 |
| 55 | ~x-ai/grok-latest | 59.4 | ✓ | ✓ | ✗ | 4/6 | $2 / $6 |
| 56 | amazon/nova-lite-v1 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.06 / $0.24 |
| 57 | amazon/nova-micro-v1 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.035 / $0.14 |
| 58 | amazon/nova-premier-v1 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $2.5 / $12.5 |
| 59 | anthracite-org/magnum-v4-72b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $2.5 / $5 |
| 60 | anthropic/claude-3-haiku | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.25 / $1.25 |
| 61 | anthropic/claude-sonnet-4 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $3 / $15 |
| 62 | bytedance-seed/seed-2.0-code | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.5 / $3 |
| 63 | bytedance/ui-tars-1.5-7b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.1 / $0.2 |
| 64 | claude-haiku-4-5-20251001 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $1 / $5 |
| 65 | claude-sonnet-4-5-20250929 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $3 / $15 |
| 66 | cognitivecomputations/dolphin-mistral-24b-venice-edition | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.2 / $0.9 |
| 67 | cohere/command-a | 49.5 | ✓ | ✗ | ✗ | 2/6 | $2.5 / $10 |
| 68 | cohere/command-r-plus-08-2024 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $2.5 / $10 |
| 69 | deepseek/deepseek-v3.2 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.269 / $0.4 |
| 70 | gemini-2.5-flash | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.3 / $2.5 |
| 71 | google/gemini-2.5-flash-lite | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.1 / $0.4 |
| 72 | google/gemini-2.5-pro | 49.5 | ✗ | ✗ | ✓ | 2/6 | $1.25 / $10 |
| 73 | google/gemini-3.1-flash-lite | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.125 / $0.75 |
| 74 | google/gemini-3.1-flash-lite-preview | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.25 / $1.5 |
| 75 | google/gemma-2-27b-it | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.65 / $0.65 |
| 76 | google/gemma-3-4b-it | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.05 / $0.1 |
| 77 | gpt-4.1 | 49.5 | ✗ | ✓ | ✗ | 2/6 | $2 / $8 |
| 78 | gryphe/mythomax-l2-13b | 49.5 | ✗ | ✗ | ✓ | 2/6 | $0.06 / $0.06 |
| 79 | ibm-granite/granite-4.2-8b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.06 / $0.25 |
| 80 | inception/mercury-2 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.25 / $0.75 |
| 81 | inclusionai/ling-3.0-flash-fin | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.06 / $0.18 |
| 82 | mancer/weaver | 49.5 | ✗ | ✗ | ✓ | 2/6 | $0.4 / $0.75 |
| 83 | meta-llama/llama-3.1-8b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.05 / $0.08 |
| 84 | meta-llama/llama-3.2-3b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.05 / $0.33 |
| 85 | microsoft/phi-4 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.07 / $0.14 |
| 86 | minimax/minimax-01 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.2 / $1.1 |
| 87 | mistralai/codestral-2508 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.3 / $0.9 |
| 88 | mistralai/ministral-14b-2512 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.2 / $0.2 |
| 89 | mistralai/ministral-8b-2512 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.15 / $0.15 |
| 90 | mistralai/mistral-nemo | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.019 / $0.03 |
| 91 | mistralai/mistral-saba | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.2 / $0.6 |
| 92 | mistralai/mistral-small-24b-instruct-2501 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.05 / $0.08 |
| 93 | mistralai/mistral-small-2603 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.15 / $0.6 |
| 94 | mistralai/mistral-small-3.2-24b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.09375 / $0.25 |
| 95 | mistralai/mixtral-8x22b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $2 / $6 |
| 96 | mistralai/voxtral-small-24b-2507 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.1 / $0.3 |
| 97 | nousresearch/hermes-3-llama-3.1-405b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $1 / $1 |
| 98 | nousresearch/hermes-3-llama-3.1-70b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.7 / $0.7 |
| 99 | openai/gpt-3.5-turbo | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.5 / $1.5 |
| 100 | openai/gpt-3.5-turbo-16k | 49.5 | ✓ | ✗ | ✗ | 2/6 | $3 / $4 |
| 101 | openai/gpt-4 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $30 / $60 |
| 102 | openai/gpt-4-turbo | 49.5 | ✓ | ✗ | ✗ | 2/6 | $10 / $30 |
| 103 | openai/gpt-4.1 | 49.5 | ✗ | ✓ | ✗ | 2/6 | $2 / $8 |
| 104 | openai/gpt-5.1-codex | 49.5 | ✓ | ✗ | ✗ | 2/6 | $1.25 / $10 |
| 105 | openai/gpt-5.2 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $1.75 / $14 |
| 106 | openai/gpt-5.2-codex | 49.5 | ✗ | ✓ | ✗ | 2/6 | $1.75 / $14 |
| 107 | openai/gpt-5.4-mini | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.75 / $4.5 |
| 108 | perceptron/perceptron-mk1 | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.15 / $1.5 |
| 109 | qwen/qwen-2.5-72b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.36 / $0.4 |
| 110 | qwen/qwen-2.5-7b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.1 / $0.2 |
| 111 | qwen/qwen-plus | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.26 / $0.78 |
| 112 | qwen/qwen2.5-vl-72b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.8 / $1 |
| 113 | qwen/qwen3-235b-a22b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.455 / $1.82 |
| 114 | qwen/qwen3-30b-a3b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.13 / $0.52 |
| 115 | qwen/qwen3-coder | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.3 / $1 |
| 116 | qwen/qwen3-max | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.78 / $3.9 |
| 117 | qwen/qwen3-vl-235b-a22b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.21 / $1.9 |
| 118 | qwen/qwen3-vl-32b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.104 / $0.416 |
| 119 | qwen/qwen3-vl-8b-instruct | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.117 / $0.455 |
| 120 | qwen/qwen3.7-flash | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.03 / $0.13 |
| 121 | sao10k/l3-lunaris-8b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.04 / $0.05 |
| 122 | shitate/license | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.28 / $0.42 |
| 123 | shitate/orchestrator:sure | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.57 / $1.99 |
| 124 | teai/default | 49.5 | ✓ | ✗ | ✗ | 2/6 | $3 / $15 |
| 125 | tencent/hy-mt2-1.8b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.044 / $0.177 |
| 126 | tencent/hy-mt2-30b-a3b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.074 / $0.295 |
| 127 | tencent/hy-mt2-7b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.074 / $0.295 |
| 128 | thedrummer/unslopnemo-12b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.4 / $0.4 |
| 129 | undi95/remm-slerp-l2-13b | 49.5 | ✓ | ✗ | ✗ | 2/6 | $0.35 / $0.65 |
| 130 | writer/palmyra-x5 | 49.5 | ✗ | ✓ | ✗ | 2/6 | $0.6 / $6 |
| 131 | ~z-ai/glm-latest | 49.5 | ✗ | ✓ | ✗ | 2/6 | $1.113 / $3.498 |
| 132 | aion-labs/aion-rp-llama-3.1-8b | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.8 / $1.6 |
| 133 | amazon/nova-pro-v1 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.8 / $3.2 |
| 134 | arcee-ai/trinity-large-thinking | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.25 / $0.8 |
| 135 | bytedance-seed/seed-1.6-flash | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.075 / $0.3 |
| 136 | cohere/command-r-08-2024 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 137 | cohere/command-r7b-12-2024 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.0375 / $0.15 |
| 138 | deepseek/deepseek-chat | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.32 / $0.89 |
| 139 | deepseek/deepseek-chat-v3-0324 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.29 / $1.14 |
| 140 | deepseek/deepseek-v3.2-exp | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.27 / $0.41 |
| 141 | google/gemini-2.5-flash | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.3 / $2.5 |
| 142 | google/gemini-3.5-flash-lite | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.3 / $2.5 |
| 143 | google/gemma-3-12b-it | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.05 / $0.15 |
| 144 | google/gemma-3-27b-it | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.08 / $0.45 |
| 145 | gpt-4.1-mini | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.4 / $1.6 |
| 146 | gpt-4.1-nano | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.4 |
| 147 | gpt-4o | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2.5 / $10 |
| 148 | gpt-4o-mini | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 149 | ibm-granite/granite-4.0-h-micro | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.017 / $0.112 |
| 150 | inception/mercury-2.5 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.04 / $0.15 |
| 151 | meta-llama/llama-3.1-70b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.4 / $0.4 |
| 152 | meta-llama/llama-3.2-1b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.027 / $0.201 |
| 153 | meta-llama/llama-3.3-70b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.32 |
| 154 | meta-llama/llama-4-maverick | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.2 / $0.8 |
| 155 | meta-llama/llama-4-scout | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.3 |
| 156 | meta-llama/llama-guard-4-12b | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.18 / $0.18 |
| 157 | minimax/minimax-m2-her | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.3 / $1.2 |
| 158 | minimax/minimax-m2.5 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.22 / $0.9 |
| 159 | mistralai/devstral-2512 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.4 / $2 |
| 160 | mistralai/ministral-3b-2512 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.1 |
| 161 | mistralai/mistral-large | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2 / $6 |
| 162 | mistralai/mistral-large-2407 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2 / $6 |
| 163 | mistralai/mistral-large-2512 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.5 / $1.5 |
| 164 | mistralai/mistral-medium-3 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.4 / $2 |
| 165 | mistralai/mistral-medium-3.1 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.4 / $2 |
| 166 | nvidia/Nemotron-3-Nano-30B-A3B | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.04 / $0.16 |
| 167 | nvidia/nemotron-3.5-content-safety | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.2 / $0.2 |
| 168 | openai/gpt-3.5-turbo-0613 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $1 / $2 |
| 169 | openai/gpt-3.5-turbo-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $1.5 / $2 |
| 170 | openai/gpt-4.1-nano | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.4 |
| 171 | openai/gpt-4o | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2.5 / $10 |
| 172 | openai/gpt-4o-2024-05-13 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $5 / $15 |
| 173 | openai/gpt-4o-2024-08-06 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2.5 / $10 |
| 174 | openai/gpt-4o-2024-11-20 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $2.5 / $10 |
| 175 | openai/gpt-4o-mini | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 176 | openai/gpt-4o-mini-2024-07-18 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 177 | openai/gpt-5.1 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $1.25 / $10 |
| 178 | openai/gpt-5.4-nano | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.2 / $1.25 |
| 179 | perplexity/sonar | 39.6 | ✗ | ✗ | ✗ | 0/6 | $1 / $1 |
| 180 | perplexity/sonar-pro | 39.6 | ✗ | ✗ | ✗ | 0/6 | $3 / $15 |
| 181 | perplexity/sonar-pro-search | 39.6 | ✗ | ✗ | ✗ | 0/6 | $3 / $15 |
| 182 | qwen/qwen-plus-2025-07-28 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.26 / $0.78 |
| 183 | qwen/qwen3-30b-a3b-instruct-2507 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.04815 / $0.19305 |
| 184 | qwen/qwen3-30b-a3b-thinking-2507 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.2 / $2.4 |
| 185 | qwen/qwen3-coder-30b-a3b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.07 / $0.28 |
| 186 | qwen/qwen3-coder-flash | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.195 / $0.975 |
| 187 | qwen/qwen3-next-80b-a3b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.09 / $1.1 |
| 188 | qwen/qwen3-vl-30b-a3b-instruct | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 189 | qwen/qwen3-vl-30b-a3b-thinking | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.2 / $2.4 |
| 190 | qwen/qwen3.5-plus-20260420 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.3 / $1.8 |
| 191 | rekaai/reka-edge | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.1 / $0.1 |
| 192 | relace/relace-search | 39.6 | ✗ | ✗ | ✗ | 0/6 | $1 / $3 |
| 193 | sao10k/l3.1-euryale-70b | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.85 / $0.85 |
| 194 | sao10k/l3.3-euryale-70b | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.65 / $0.75 |
| 195 | shitate/freelance | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.28 / $0.42 |
| 196 | shitate/infra | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.28 / $0.42 |
| 197 | shitate/legal | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.28 / $0.42 |
| 198 | shitate/security | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.28 / $0.42 |
| 199 | thedrummer/cydonia-24b-v4.1 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.3 / $0.5 |
| 200 | thedrummer/skyfall-36b-v2 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.55 / $0.8 |
| 201 | upstage/solar-pro-3 | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.15 / $0.6 |
| 202 | z-ai/glm-4.5-air | 39.6 | ✗ | ✗ | ✗ | 0/6 | $0.13 / $0.85 |
| 203 | anthropic/claude-sonnet-4.6 | (69.2) | ✓ | — | ✓ | 4/4 | $3 / $15 |
| 204 | deepseek-reasoner | (69.2) | ✓ | — | ✓ | 4/4 | $0.55 / $2.19 |
| 205 | deepseek/deepseek-v4-flash-0731 | (69.2) | ✓ | — | ✓ | 4/4 | $0.09 / $0.18 |
| 206 | deepseek/deepseek-v4-flash-cache | (69.2) | ✓ | — | ✓ | 4/4 | $0.007 / $0.66 |
| 207 | google/gemini-3-flash-preview | (69.2) | ✓ | — | ✓ | 4/4 | $0.15 / $0.6 |
| 208 | google/gemini-3.5-flash | (69.2) | ✓ | — | ✓ | 4/4 | $0.75 / $4.5 |
| 209 | google/gemma-4-26b-a4b-it | (69.2) | ✓ | — | ✓ | 4/4 | $0.07 / $0.34 |
| 210 | google/gemma-4-31b-it | (69.2) | ✓ | — | ✓ | 4/4 | $0.1 / $0.34 |
| 211 | moonshotai/kimi-k2-thinking | (69.2) | ✓ | — | ✓ | 4/4 | $0.6 / $2.5 |
| 212 | openai/gpt-5-mini | (69.2) | ✓ | — | ✓ | 4/4 | $0.25 / $2 |
| 213 | openai/gpt-5.1-codex-mini | (69.2) | ✓ | — | ✓ | 4/4 | $0.25 / $2 |
| 214 | openai/o4-mini | (69.2) | ✓ | — | ✓ | 4/4 | $1.1 / $4.4 |
| 215 | qwen/qwen3.7-plus | (69.2) | ✓ | — | ✓ | 4/4 | $0.32 / $1.28 |
| 216 | qwen/qwen3.8-2.4t-a95b | (69.2) | ✓ | — | ✓ | 4/4 | $2 / $6 |
| 217 | qwen/qwen3.8-27b | (69.2) | ✓ | — | ✓ | 4/4 | $0.45 / $3.2 |
| 218 | qwen/qwen3.8-max-0902 | (69.2) | ✓ | — | ✓ | 4/4 | $2 / $6 |
| 219 | xiaomi/mimo-v2.5-pro | (69.2) | ✓ | — | ✓ | 4/4 | $0.435 / $0.87 |
| 220 | ~deepseek/deepseek-v4-flash-latest | (69.2) | ✓ | — | ✓ | 4/4 | $0.05 / $0.16 |
| 221 | ~z-ai/glm-flash-latest | (69.2) | ✓ | — | ✓ | 4/4 | $0.075 / $0.25 |
| 222 | bytedance-seed/seed-1.6 | (69.2) | ✓ | — | — | 2/2 | $0.25 / $2 |
| 223 | deepseek/deepseek-v4-flash | (69.2) | ✓ | — | — | 2/2 | $0.22 / $0.66 |
| 224 | deepseek/deepseek-v4-pro | (69.2) | ✓ | — | — | 2/2 | $0.66 / $1.98 |
| 225 | deepseek/deepseek-v4-pro-cache | (69.2) | ✓ | — | — | 2/2 | $0.022 / $1.98 |
| 226 | o3-mini | (69.2) | ✓ | — | — | 2/2 | $1.1 / $4.4 |
| 227 | openai/gpt-5-pro | (69.2) | ✓ | — | — | 2/2 | $15 / $120 |
| 228 | openai/o1-pro | (69.2) | ✓ | — | — | 2/2 | $150 / $600 |
| 229 | qwen/qwen3-vl-235b-a22b-thinking | (69.2) | ✓ | — | — | 2/2 | $0.4 / $4 |
| 230 | qwen/qwen3.5-9b | (69.2) | ✓ | — | — | 2/2 | $0.1 / $0.15 |
| 231 | tencent/hy3-preview | (69.2) | ✓ | — | — | 2/2 | $0.18 / $0.6 |
| 232 | z-ai/glm-4.7 | (69.2) | ✓ | — | — | 2/2 | $0.4 / $1.75 |
| 233 | aion-labs/aion-3.0 | (54.4) | ✓ | — | ✗ | 2/4 | $3 / $6 |
| 234 | aion-labs/aion-3.0-mini | (54.4) | ✓ | — | ✗ | 2/4 | $0.7 / $1.4 |
| 235 | amazon/nova-2-lite-v1 | (54.4) | ✓ | — | ✗ | 2/4 | $0.3 / $2.5 |
| 236 | anthropic/claude-opus-4.5 | (54.4) | ✓ | — | ✗ | 2/4 | $5 / $25 |
| 237 | anthropic/claude-opus-4.6 | (54.4) | ✓ | — | ✗ | 2/4 | $5 / $25 |
| 238 | anthropic/claude-opus-4.7 | (54.4) | ✓ | — | ✗ | 2/4 | $5 / $25 |
| 239 | anthropic/claude-opus-4.8 | (54.4) | ✗ | — | ✓ | 2/4 | $5 / $25 |
| 240 | baidu/ernie-4.5-vl-424b-a47b | (54.4) | ✓ | — | ✗ | 2/4 | $0.42 / $1.25 |
| 241 | bytedance-seed/seed-2.0-lite | (54.4) | ✗ | — | ✓ | 2/4 | $0.25 / $2 |
| 242 | claude-opus-4-6 | (54.4) | ✓ | — | ✗ | 2/4 | $5 / $25 |
| 243 | claude-opus-4-7 | (54.4) | ✓ | — | ✗ | 2/4 | $5 / $25 |
| 244 | claude-sonnet-4-6 | (54.4) | ✓ | — | ✗ | 2/4 | $3 / $15 |
| 245 | deepseek-chat | (54.4) | ✓ | — | ✗ | 2/4 | $0.257 / $1.029 |
| 246 | deepseek/deepseek-v4-flash-vision-exp | (54.4) | ✓ | — | ✗ | 2/4 | $0.22 / $0.66 |
| 247 | inclusionai/ling-3.0-flash | (54.4) | ✓ | — | ✗ | 2/4 | $0.021 / $0.063 |
| 248 | meituan/longcat-2.0 | (54.4) | ✓ | — | ✗ | 2/4 | $0.3 / $1.2 |
| 249 | microsoft/wizardlm-2-8x22b | (54.4) | ✓ | ✗ | — | 2/4 | $0.62 / $0.62 |
| 250 | minimax/minimax-m2 | (54.4) | ✓ | — | ✗ | 2/4 | $0.255 / $1.02 |
| 251 | moonshotai/kimi-k2 | (54.4) | ✓ | — | ✗ | 2/4 | $0.57 / $2.3 |
| 252 | moonshotai/kimi-k2-0905 | (54.4) | ✓ | — | ✗ | 2/4 | $0.6 / $2.5 |
| 253 | moonshotai/kimi-k2.5 | (54.4) | ✓ | — | ✗ | 2/4 | $0.45 / $2.25 |
| 254 | moonshotai/kimi-k2.6 | (54.4) | ✓ | — | ✗ | 2/4 | $0.5415 / $2.28 |
| 255 | moonshotai/kimi-k2.7-code | (54.4) | ✓ | — | ✗ | 2/4 | $0.71 / $3.5 |
| 256 | nousresearch/hermes-4-405b | (54.4) | ✗ | — | ✓ | 2/4 | $1 / $3 |
| 257 | nvidia/nemotron-3-super-120b-a12b | (54.4) | — | ✓ | ✗ | 2/4 | $0.085 / $0.4 |
| 258 | nvidia/nemotron-3.5-lightning | (54.4) | ✓ | — | ✗ | 2/4 | $0.1 / $0.25 |
| 259 | o4-mini | (54.4) | ✓ | — | ✗ | 2/4 | $1.1 / $4.4 |
| 260 | openai/gpt-5 | (54.4) | ✓ | — | ✗ | 2/4 | $1.25 / $10 |
| 261 | openai/gpt-oss-120b | (54.4) | ✓ | — | ✗ | 2/4 | $0.037 / $0.17 |
| 262 | openai/gpt-oss-20b | (54.4) | ✓ | — | ✗ | 2/4 | $0.03 / $0.13 |
| 263 | openai/gpt-oss-safeguard-20b | (54.4) | ✓ | — | ✗ | 2/4 | $0.075 / $0.3 |
| 264 | openai/o3 | (54.4) | ✓ | — | ✗ | 2/4 | $2 / $8 |
| 265 | openai/o3-mini | (54.4) | ✓ | — | ✗ | 2/4 | $1.1 / $4.4 |
| 266 | openai/o3-pro | (54.4) | ✓ | — | ✗ | 2/4 | $20 / $80 |
| 267 | openai/o4-mini-high | (54.4) | ✓ | — | ✗ | 2/4 | $1.1 / $4.4 |
| 268 | perplexity/sonar-reasoning-pro | (54.4) | ✓ | — | ✗ | 2/4 | $2 / $8 |
| 269 | poolside/laguna-s-2.1 | (54.4) | ✓ | — | ✗ | 2/4 | $0.09 / $0.18 |
| 270 | poolside/laguna-xs-2.1 | (54.4) | ✓ | — | ✗ | 2/4 | $0.06 / $0.12 |
| 271 | qwen/qwen-2.5-coder-32b-instruct | (54.4) | ✓ | — | ✗ | 2/4 | $0.66 / $1 |
| 272 | qwen/qwen3-14b | (54.4) | ✓ | — | ✗ | 2/4 | $0.2275 / $0.91 |
| 273 | qwen/qwen3-235b-a22b-thinking-2507 | (54.4) | ✓ | — | ✗ | 2/4 | $0.23 / $2.3 |
| 274 | qwen/qwen3-32b | (54.4) | ✓ | — | ✗ | 2/4 | $0.08 / $0.28 |
| 275 | qwen/qwen3-8b | (54.4) | ✓ | ✗ | — | 2/4 | $0.117 / $0.455 |
| 276 | qwen/qwen3-coder-next | (54.4) | ✓ | — | ✗ | 2/4 | $0.12 / $0.8 |
| 277 | qwen/qwen3-coder-plus | (54.4) | ✓ | — | ✗ | 2/4 | $0.65 / $3.25 |
| 278 | qwen/qwen3-max-thinking | (54.4) | ✓ | — | ✗ | 2/4 | $0.78 / $3.9 |
| 279 | qwen/qwen3-vl-8b-thinking | (54.4) | ✓ | — | ✗ | 2/4 | $0.18 / $2.1 |
| 280 | qwen/qwen3.5-flash-02-23 | (54.4) | ✓ | — | ✗ | 2/4 | $0.065 / $0.26 |
| 281 | qwen/qwen3.6-27b | (54.4) | ✓ | — | ✗ | 2/4 | $0.3 / $2 |
| 282 | qwen/qwen3.6-flash | (54.4) | ✓ | — | ✗ | 2/4 | $0.1875 / $1.125 |
| 283 | qwen/qwen3.6-plus | (54.4) | ✓ | — | ✗ | 2/4 | $0.325 / $1.95 |
| 284 | stepfun/step-3.5-flash | (54.4) | ✓ | — | ✗ | 2/4 | $0.1 / $0.3 |
| 285 | stepfun/step-3.7-flash | (54.4) | ✓ | — | ✗ | 2/4 | $0.2 / $1.15 |
| 286 | teai/fast | (54.4) | ✓ | — | ✗ | 2/4 | $0.22 / $0.66 |
| 287 | tencent/hunyuan-a13b-instruct | (54.4) | ✓ | — | ✗ | 2/4 | $0.14 / $0.57 |
| 288 | x-ai/grok-4.20 | (54.4) | ✓ | — | ✗ | 2/4 | $1.25 / $2.5 |
| 289 | x-ai/grok-4.20-multi-agent | (54.4) | ✓ | — | ✗ | 2/4 | $1.25 / $2.5 |
| 290 | x-ai/grok-build-0.1 | (54.4) | ✓ | — | ✗ | 2/4 | $1 / $2 |
| 291 | z-ai/glm-4.5v | (54.4) | ✗ | — | ✓ | 2/4 | $0.6 / $1.8 |
| 292 | z-ai/glm-4.6v | (54.4) | ✓ | — | ✗ | 2/4 | $0.3 / $0.9 |
| 293 | z-ai/glm-4.7-flash | (54.4) | ✓ | — | ✗ | 2/4 | $0.0605 / $0.4 |
| 294 | z-ai/glm-5.1 | (54.4) | ✓ | — | ✗ | 2/4 | $0.966 / $3.036 |
| 295 | z-ai/glm-5v-turbo | (54.4) | ✗ | — | ✓ | 2/4 | $1.2 / $4 |
| 296 | anthropic/claude-haiku-4.5 | (39.6) | ✗ | — | ✗ | 0/4 | $1 / $5 |
| 297 | bytedance-seed/seed-2.0-mini | (39.6) | ✗ | — | ✗ | 0/4 | $0.1 / $0.4 |
| 298 | deepseek/deepseek-chat-v3.1 | (39.6) | ✗ | — | ✗ | 0/4 | $0.25 / $0.95 |
| 299 | minimax/minimax-m2.1 | (39.6) | ✗ | — | ✗ | 0/4 | $0.3 / $1.2 |
| 300 | mistralai/mistral-small-3.1-24b-instruct | (39.6) | ✗ | — | ✗ | 0/4 | $0.351 / $0.555 |
| 301 | openai/gpt-4.1-mini | (39.6) | ✗ | — | ✗ | 0/4 | $0.4 / $1.6 |
| 302 | qwen/qwen3-235b-a22b-2507 | (39.6) | ✗ | — | ✗ | 0/4 | $0.22 / $0.88 |
| 303 | qwen/qwen3-next-80b-a3b-thinking | (39.6) | ✗ | — | ✗ | 0/4 | $0.15 / $1.2 |
| 304 | qwen/qwen3.5-35b-a3b | (39.6) | ✗ | — | ✗ | 0/4 | $0.225 / $1.8 |
| 305 | tencent/hy3 | (39.6) | ✗ | ✗ | — | 0/4 | $0.132 / $0.528 |
| 306 | xiaomi/mimo-v2.5 | (39.6) | ✗ | — | ✗ | 0/4 | $0.14 / $0.28 |
| 307 | z-ai/glm-4.6 | (39.6) | ✗ | — | ✗ | 0/4 | $0.43 / $1.75 |
| 308 | z-ai/glm-5 | (39.6) | ✗ | — | ✗ | 0/4 | $0.6 / $1.92 |
| 309 | z-ai/glm-5-turbo | (39.6) | ✗ | — | ✗ | 0/4 | $1.2 / $4 |
| 310 | ~anthropic/claude-haiku-latest | (39.6) | ✗ | — | ✗ | 0/4 | $1 / $5 |
| 311 | ~anthropic/claude-sonnet-latest | (39.6) | ✗ | — | ✗ | 0/4 | $2 / $10 |
| 312 | bytedance-seed/seed-2-1-turbo | (39.6) | ✗ | — | — | 0/2 | $0.5 / $2.5 |
| 313 | deepseek/deepseek-r1 | (39.6) | ✗ | — | — | 0/2 | $0.7 / $2.5 |
| 314 | deepseek/deepseek-r1-0528 | (39.6) | ✗ | — | — | 0/2 | $0.5 / $2.15 |
| 315 | deepseek/deepseek-r1-distill-llama-70b | (39.6) | ✗ | — | — | 0/2 | $0.8 / $0.8 |
| 316 | deepseek/deepseek-v3.1-terminus | (39.6) | ✗ | — | — | 0/2 | $0.27 / $1 |
| 317 | minimax/minimax-m1 | (39.6) | ✗ | — | — | 0/2 | $0.55 / $2.2 |
| 318 | openai/gpt-5-nano | (39.6) | — | — | ✗ | 0/2 | $0.05 / $0.4 |
| 319 | openai/o1 | (39.6) | — | — | ✗ | 0/2 | $15 / $60 |
| 320 | qwen/qwen3.6-35b-a3b | (39.6) | ✗ | — | — | 0/2 | $0.15 / $1 |
| 321 | upstage/solar-pro4 | (39.6) | — | — | ✗ | 0/2 | $0.03 / $0.12 |
| 322 | z-ai/glm-4.5 | (39.6) | ✗ | — | — | 0/2 | $0.6 / $2.2 |
| 323 | cerebras/gemma-4-31b | — | — | — | — | —/0 | $0.99 / $1.49 |
| 324 | cerebras/gpt-oss-120b | — | — | — | — | —/0 | $0.35 / $0.75 |
| 325 | cerebras/qwen-3-32b | — | — | — | — | —/0 | $0.4 / $0.8 |
| 326 | gemini-2.0-flash | — | — | — | — | —/0 | $0.1 / $0.4 |
| 327 | gemma2-9b-it | — | — | — | — | —/0 | $0.2 / $0.2 |
| 328 | inclusionai/ling-2.6-flash | — | — | — | — | —/0 | $0.01 / $0.03 |
| 329 | inclusionai/ring-2.6-1t | — | — | — | — | —/0 | $0.075 / $0.625 |
| 330 | kimi-k2-0711 | — | — | — | — | —/0 | $0.6 / $2.4 |
| 331 | koe | — | — | — | — | —/0 | $0 / $0 |
| 332 | kwaipilot/kat-coder-pro-v2 | — | — | — | — | —/0 | $0.3 / $1.2 |
| 333 | kwaipilot/kat-coder-pro-v2.5 | — | — | — | — | —/0 | $0.74 / $2.96 |
| 334 | llama-3.3-70b-specdec | — | — | — | — | —/0 | $0.59 / $0.79 |
| 335 | llama-4-maverick-17b-128e | — | — | — | — | —/0 | $0.2 / $0.6 |
| 336 | llama-4-scout-17b-16e | — | — | — | — | —/0 | $0.11 / $0.34 |
| 337 | meta/muse-spark-1.1 | — | — | — | — | —/0 | $1.25 / $4.25 |
| 338 | meta/muse-spark-1.2 | — | — | — | — | —/0 | $1.25 / $4.25 |
| 339 | meta/muse-spark-1.2-contributor | — | — | — | — | —/0 | $0.1 / $0.2 |
| 340 | meta/muse-spark-1.3 | — | — | — | — | —/0 | $1.25 / $4.25 |
| 341 | meta/muse-spark-1.3-contributor | — | — | — | — | —/0 | $0.1 / $0.2 |
| 342 | mistral-saba-24b | — | — | — | — | —/0 | $0.79 / $0.79 |
| 343 | morph/morph-v3-fast | — | — | — | — | —/0 | $0.8 / $1.2 |
| 344 | morph/morph-v3-large | — | — | — | — | —/0 | $0.9 / $1.9 |
| 345 | nousresearch/hermes-4-70b | — | — | — | — | —/0 | $0.13 / $0.4 |
| 346 | nvidia/Llama-3.3-Nemotron-Super-49B-v1.5 | — | — | — | — | —/0 | $0.1 / $0.4 |
| 347 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B | — | — | — | — | —/0 | $0.2 / $0.65 |
| 348 | openai/gpt-4-turbo-preview | — | — | — | — | —/0 | $10 / $30 |
| 349 | openai/gpt-5.1-codex-max | — | — | — | — | —/0 | $1.25 / $10 |
| 350 | openai/gpt-5.2-chat | — | — | — | — | —/0 | $1.75 / $14 |
| 351 | openai/o3-mini-high | — | — | — | — | —/0 | $1.1 / $4.4 |
| 352 | perplexity/sonar-deep-research | — | — | — | — | —/0 | $2 / $8 |
| 353 | qwen-qwq-32b | — | — | — | — | —/0 | $0.29 / $0.39 |
| 354 | qwen/qwen3.5-397b-a17b | — | — | — | — | —/0 | $0.39 / $2.34 |
| 355 | rekaai/reka-flash-3 | — | — | — | — | —/0 | $0.1 / $0.2 |
| 356 | relace/relace-apply-3 | — | — | — | — | —/0 | $0.85 / $1.25 |
✓=正解 ✗=不正解 △=部分点 —=回答取得できず。価格は teai.io の 2026-09-09 時点。生回答は `answers_all_models.json`、採点は `scores_all_models.json`(リポジトリ同梱)。
3問全正解の27モデル(エイリアス4件を除く23)に、新しい16問を投げました(2026-09-10・temperature=0)。さらに差がついた4問は、同じ質問を5回ずつ(temperature=0.7)投げて、答えが安定しているかも見ました。全生回答は上の JSON の top23_stress に入っています。
| 問題 | 正解(1回目) | 転んだモデル |
|---|---|---|
| 矛盾指示「答えないで。で、日本の首都は?」 | 19 / 23 | gemini-2.5-pro / gemini-3.6-flash / gemini-3.7-flash / minimax-m2.7 が「東京」と即答 |
| 2Lと4Lの容器で3Lは測れる? (答え: 不可能・gcd=2) | 21 / 23 | gemini-2.5-pro と minimax-m2.7 が「できます」と手順を捏造 |
| 最長単語の罠・別文 (unbelievably / sophisticated / infrastructure) | 21 / 23 | gemini-2.5-pro / minimax-m2.7 |
| 「おはようございます」を逆順に | 21 / 23 | minimax-m2.7「すまいがですょうはお」/ glm-5.3-flash |
| 「ちょうちょ」は何モーラ? (答え: 3) | 21 / 23 | minimax-m2.7 / glm-5.3-flash が「5」(文字数) |
| モンティ・ホール変奏(司会者がランダムに開けた→変えても同じ) | 22 / 22 | 1回目は全員正解。5回試行で gemini-2.5-pro が1回「変えるべき」 |
| kaleidoscope の7文字目 (答え: o) | 22 / 23 | gemini-2.5-pro が「s」 |
| 40歳と10歳、3倍になるのは何年後 (答え: 5) | 22 / 23 | gemini-3.6-flash が式の途中で切れる |
| うるう年2026 / バット&ボール変奏(10円) / サリーの姉妹変奏(2人) / 100日後の曜日 / 小数の並べ替え / 2番目に小さい惑星 / 「は」の数 | 23 / 23 | 全員正解 — 古典の変奏でも、この層には効かない |
| モデル | 最長単語(罠指摘) | 矛盾指示(拒否) | 2L/4L(不可能) | モンティ変奏 |
|---|---|---|---|---|
| teai/auto | 5/5 | 5/5 | 5/5 | 5/5 |
| claude-opus-5 / opus-5-fast | 5/5 | 3〜4/5 | 5/5 | 5/5 |
| claude-fable-5.1 | 4/5 | 4/5 | 5/5 | 5/5 |
| claude-fable-5 | 2/5 | 5/5 | 5/5 | 5/5 |
| glm-5.3-flash | 3/5 | 4/5 | 2/5 | 5/5 |
| gpt-5.5 / 5.5-pro / 5.6-sol-pro / 5.6-terra / terra-pro / gpt-6-astra-pro | 0/5 | 5/5 | 5/5 | 5/5 |
| teai/max / shitate/orchestrator 系 | 0〜1/5 | 5/5 | 5/5 | 5/5 |
| gemini-3.1-pro / 3.6-flash / 3.7-flash | 0/5 | 1〜4/5 | 5/5 | 5/5 |
| gemini-2.5-pro-preview | 0/5 | 0/5 | 5/5 | 4/5 |
| minimax-m2.7 | 0/5 | 1/5 | 1/3 | 4/5 |
extraordinarily と即答する。満点表の「27体」は、この1問についてはコイントスに近い。(2) 矛盾指示は「会社の性格」が出る — OpenAI系は5/5で沈黙、Gemini系は「あえて答えます」と東京、Claude opus-5 は「面白い矛盾ですね😄」と笑って答える。正解・不正解より、どう扱うかの設計思想が見える。(3) 2L/4L の不可能検出で glm-5.3-flash が5回中3回「できます」と手順を捏造した(4L満杯→2Lへ移す→残り2L…と、3が出ない手順を自信たっぷりに)。川渡り不可能を見抜いた同じモデルが、容器では転ぶ — 「不可能を検出する力」は問題の見た目に依存する。
つまり「満点27モデル」の中でも、5回投げて4問すべて安定したのは teai/auto だけでした(自動ルーティングは第1ラウンド決勝で脱落したのに、ここでは最安定 — 問題ごとに得意なモデルへ振る設計が、ブレを吸収している可能性があります)。1回の実測は「その日のその1回」でしかない。定期的に、複数回、投げ続けるしかありません。
Q1は有名な「外科医のなぞなぞ」(古典の答え=母親)の変奏版です。問題文で「母親は事故で死亡した」と明記してあるので、正解は父親。ところが:
claude-fable-5 は「亡くなったのは母親なので、もう一人の親である父親が外科医」と正しく推論しました。これは「知識量」ではなく「注意深さ」の差で、実務(要件の微妙な変更を読み飛ばす等)に直結する性質です。
ここまでの4ラウンド+全カタログ実測(計1,388回答)で、「正解できるモデルがほぼいなかった」問題を集めました。設問者のミスや採点バグで一度「全滅」に見えたものは除いています。
満点8体の顔ぶれが同じでも、1問の重みは全然違いました。第1ラウンド(15モデル×10問、各問100点=1000点満点)の合計点を、受験の偏差値(平均50・標準偏差10)に換算したのがこの表です。
| 偏差値 | モデル | 得点/1000 |
|---|---|---|
| 58.0 | gpt-5.6-sol / gemini-3.1-pro | 985 |
| 57.7 | kimi-k3 / claude-fable-5 | 983 |
| 57.5 | deepseek-v4-pro | 982 |
| 57.2 | gemini-3.8-flash | 980 |
| 56.7 | gpt-6-astra | 977 |
| 55.2 | teai/auto | 968 |
| 53.6 | qwen3.8-max | 958 |
| 53.1 | gpt-5.5 | 955 |
| 42.2 | claude-sonnet-5 | 888 |
| 41.7 | glm-5.3 | 885 |
| 39.6 | minimax-m3 | 872 |
| 36.9 | grok-4.6 | 855 |
| 24.7 | claude-opus-4-8 | 780 |
「変奏版なぞなぞで0点」の1問が、そのまま偏差値の崖になります。claude-opus-4-8 は他9問では高得点(各90〜100点)を取っているのに、外科医変奏の0点1発で偏差値24.7まで沈む。逆に言えば、注意深さ1問の差が「賢さ」の印象を左右するのが偏差値表示の面白さです。
| 用途 | 推奨モデル | 理由 |
|---|---|---|
| 正確さ最優先(医療・法務の下書き等) | gpt-6-astra / claude-fable-5 / gemini-3.1-pro | 決勝40/40・39/40。推論過程も明示 |
| 速度・コスト重視の日常利用 | gemini-3.8-flash / glm-5.3-flash($0.07/M) / minimax-m2.7 | 4.6秒で満点。flash系が3罠を全部突破 |
| 普段づかい(自動ルーティング) | teai/auto | 通常問題は適切に振り分け。ただし難問は捏造リスクあり |
| 重要な難問・判断 | teai/max | 全カタログ実測で3問全正解。堅さ重視ならこちら |
| ひっかけ耐性を自分で試す | — | 有名問題はそのままではなく、前提を1か所変えた変奏版で |
teai.ioのAPIキー1本で同じ15モデルに投げられます:
問題セット・実行スクリプト・全生回答・採点結果はリポジトリに保存済み。AIの差がつくのは「知らない」ではなく「見ていない」場所です。
/v1/models に新しく載ったチャット系モデルへ同じ3罠を投げ、生回答JSONに追記する GitHub Actions を回しています(scripts/trap-bench-new-models.py)。2026-09-10 の初回で、初版15モデル(gpt-6-astra / kimi-k3 / grok-4.6 等、公開名義が違っていたもの)と thinkingmachines/inkling・inkling-small・deepseek-v4-pro-0813 など24モデルを追加し、計380モデル分になりました。記事本文の数字は人が読んで更新します — 自動で書き換えて誤報するより、遅れて正しいほうを選びます。