Сводный рейтинг Neira объединяет независимые тесты, сохраняя точные версии моделей и покрытие категорий.
Последний полный снимок: 23 авг. 2026 г., 06:02. Чат 30%, знания и рассуждение 20%, код 40%, вызов инструментов 10%.
Оценка 0–100 нормализована по позиции в каждом источнике, поэтому исходные шкалы не смешиваются.
| Ранг | Модель | Разработчик | Оценка | Покрытие | Источники |
|---|---|---|---|---|---|
#1 | moonshotai/Kimi-K3 | moonshotai | 100.0 | 1/1 | DeepSWE |
#2 | moonshotai/Kimi-K3 (Kimi Code harness.) | moonshotai | 100.0 | 1/1 | Terminal-Bench 2.1 |
Открытые веса · 4495 моделей
| O3 (High) |
| OpenAI |
| 100.0 |
| 1/1 |
| LiveCodeBench |
#4 | deepseek-ai/DeepSeek-V4-Pro-0813 (DeepSeek Harness (minimal mode), max reasoning effort, temperature=1.0, top_p=0.95.) | deepseek-ai | 96.0 | 1/1 | Terminal-Bench 2.1 |
#5 | ornith-ai/Ornith-1.0-397B | ornith-ai | 95.7 | 1/1 | SWE-bench Multilingual |
#6 | DeepSeek-R1-0528 | DeepSeek | 95.0 | 1/1 | LiveCodeBench |
#7 | zai-org/GLM-5.1 (agent: Terminus 2(Claude Code)) | zai-org | 94.7 | 1/1 | Terminal-Bench 2.0 |
#8 | deepreinforce-ai/Ornith-1.0-397B | deepreinforce-ai | 93.9 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#9 | mindlab-research/Macaron-V1-Venti | mindlab-research | 93.4 | 1/1 | DeepSWE, SWE-bench Verified |
#10 | Qwen/Qwen3.8-2.4T-A95B (Reported under the "Qwen3.8-Max" product tier (vision input, non-thinking support, 1M context by default, built-in tools) -- may not be identical to the raw open checkpoint alone. Evaluated with Claude Code (avg@10), 5h timeout, max_tokens=131072.) | qwen | 92.0 | 1/1 | Terminal-Bench 2.1 |
#11 | Gemini-2.5-Pro-06-05 | 90.0 | 1/1 | LiveCodeBench |
#12 | MiniMaxAI/MiniMax-M3 (Evaluated on internal infrastructure using Claude Code as the scaffolding. Each test was run 4 times and the average was taken.) | minimaxai | 90.0 | 1/1 | SWE-bench Verified |
#13 | ornith-ai/Ornith-1.5-397B | ornith-ai | 88.1 | 1/1 | DeepSWE, SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#14 | ornith-ai/Ornith-1.5-397B (Terminus-2 harness (Harbor framework), parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, avg of 5 runs.) | ornith-ai | 88.0 | 1/1 | Terminal-Bench 2.1 |
#15 | Qwen/Qwen3.8-27B (Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, 256K context; baseline models re-evaluated on the same refined task set.) | qwen | 87.5 | 1/1 | SWE-bench Pro |
#16 | Gemini-2.5-Pro-05-06 | 85.0 | 1/1 | LiveCodeBench |
#17 | ornith-ai/Ornith-1.5-397B (Claude Code 2.1.126 harness, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.) | ornith-ai | 84.0 | 1/1 | Terminal-Bench 2.1 |
#18 | Qwen/Qwen3.8-2.4T-A95B | qwen | 83.3 | 1/1 | DeepSWE, SWE-bench Pro |
#19 | moonshotai/Kimi-K2.6 | moonshotai | 82.5 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#20 | Qwen3-235B-A22B | Alibaba | 80.0 | 1/1 | LiveCodeBench |
#21 | zai-org/GLM-5.2 (Best reported harness: Claude Code 2.1.167 (temperature=1.0, top_p=0.95, max_new_tokens=131072, no wall-clock limit, averaged over 5 runs).) | zai-org | 80.0 | 1/1 | Terminal-Bench 2.1 |
#22 | dots-studio/dots3-note-prev | dots-studio | 78.1 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#23 | deepseek-ai/DeepSeek-V4-Flash-0731 (DeepSeek Harness (minimal mode), max reasoning effort, temperature=1.0, top_p=0.95.) | deepseek-ai | 76.0 | 1/1 | Terminal-Bench 2.1 |
#24 | XiaomiMiMo/MiMo-V2.5-Pro | xiaomimimo | 75.6 | 1/1 | SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#25 | Grok-3-Mini (High) | xAI | 75.0 | 1/1 | LiveCodeBench |
#26 | MiniMaxAI/MiniMax-M3 (Evaluated on internal infrastructure using Claude Code as the scaffolding. Testing logic is aligned with the official evaluation.) | minimaxai | 75.0 | 1/1 | SWE-bench Pro |
#27 | deepseek-ai/DeepSeek-V4-Pro | deepseek-ai | 74.1 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#28 | zai-org/GLM-5 (Z.ai reported number) | zai-org | 72.0 | 1/1 | SWE-bench Verified |
#29 | zai-org/GLM-5.2 (Terminus-2 harness (temperature=1.0, top_p=1.0, 256K context, max_episodes=500).) | zai-org | 72.0 | 1/1 | Terminal-Bench 2.1 |
#30 | mistralai/Mistral-Medium-3.5-128B | mistralai | 70.0 | 1/1 | SWE-bench Verified |
#31 | O3-Mini-2025-01-31 (High) | OpenAI | 70.0 | 1/1 | LiveCodeBench |
#32 | zai-org/GLM-5 | zai-org | 69.6 | 1/1 | SWE-bench Multilingual |
#33 | ornith-ai/Ornith-1.0-397B (Claude Code 2.1.126 harness, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.) | ornith-ai | 68.0 | 1/1 | Terminal-Bench 2.1 |
#34 | zai-org/GLM-5.1 (high reasoning) | zai-org | 65.6 | 1/1 | SWE-bench Pro |
#35 | Gemini-2.5-Flash-05-20 | 65.0 | 1/1 | LiveCodeBench |
#36 | thinkingmachines/Inkling-Small | thinkingmachines | 64.9 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#37 | deepseek-ai/DeepSeek-V4-Flash | deepseek-ai | 64.9 | 1/1 | SWE-bench Multilingual, SWE-bench Verified, Terminal-Bench 2.0 |
#38 | ornith-ai/Ornith-1.0-397B (Terminus-2 harness (Harbor framework), parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, avg of 5 runs.) | ornith-ai | 64.0 | 1/1 | Terminal-Bench 2.1 |
#39 | MiniMaxAI/MiniMax-M2.7 | minimaxai | 63.2 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, Terminal-Bench 2.0 |
#40 | zai-org/GLM-5.1 | zai-org | 63.2 | 1/1 | Terminal-Bench 2.0 |
#41 | zai-org/GLM-5.2 | zai-org | 62.0 | 1/1 | DeepSWE, SWE-bench Pro |
#42 | dots-studio/dots3-note-prev (Terminus-2 harness, parser=json, 10h timeout, temperature=0.7, top_p=0.95, max_tokens=81920, 256K context.) | dots-studio | 60.0 | 1/1 | Terminal-Bench 2.1 |
#43 | O3-Mini-2025-01-31 (Med) | OpenAI | 60.0 | 1/1 | LiveCodeBench |
#44 | Qwen/Qwen3.6-27B (Harbor/Terminus-2 harness; 3h timeout, 32 CPU/48 GB RAM; temp=1.0, top_p=0.95, top_k=20, max_tokens=80K, 256K ctx; avg of 5 runs.) | qwen | 57.9 | 1/1 | Terminal-Bench 2.0 |
#45 | XiaomiMiMo/MiMo-V2.5 | xiaomimimo | 57.6 | 1/1 | SWE-bench Pro, Terminal-Bench 2.0 |
#46 | inclusionAI/Ling-3.0-flash | inclusionai | 56.4 | 1/1 | SWE-bench Multilingual, SWE-bench Pro |
#47 | Qwen/Qwen3.8-27B (Row labeled "(Terminus)" as the harness; no further hyperparameter footnote given for this row.) | qwen | 56.0 | 1/1 | Terminal-Bench 2.1 |
#48 | Gemini-2.5-Flash-04-17 | 55.0 | 1/1 | LiveCodeBench |
#49 | papers/2603.16790 | papers | 54.0 | 1/1 | SWE-bench Verified |
#50 | tencent/Hy3 | tencent | 53.7 | 1/1 | DeepSWE, SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#51 | tencent/Hy3 (Terminus-2 scaffold, parser=xml, agent timeout=4h, max episodes=500, highest reasoning-effort tier.) | tencent | 52.0 | 1/1 | Terminal-Bench 2.1 |
#52 | thinkingmachines/Inkling | thinkingmachines | 51.2 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#53 | ornith-ai/Ornith-1.5-35B-A3B | ornith-ai | 50.6 | 1/1 | DeepSWE, SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#54 | O3-Mini-2025-01-31 (Low) | OpenAI | 50.0 | 1/1 | LiveCodeBench |
#55 | Qwen/Qwen3.6-27B | qwen | 48.4 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#56 | Qwen/Qwen3.5-397B-A17B | qwen | 48.1 | 1/1 | SWE-bench Multilingual, SWE-bench Verified, Terminal-Bench 2.0 |
#57 | meituan-longcat/LongCat-2.0 (Measured in-house under LongCat's own unified harness; no further scaffold/temperature/context details given.) | meituan-longcat | 48.0 | 1/1 | Terminal-Bench 2.1 |
#58 | MiniMaxAI/MiniMax-M2.1 | minimaxai | 46.0 | 1/1 | SWE-bench Verified |
#59 | meta-models/Muse-Glimmer-30B | meta-models | 45.1 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#60 | tencent/Hy3-preview | tencent | 45.1 | 1/1 | SWE-bench Verified, Terminal-Bench 2.0 |
#61 | Claude-Opus-4 (Thinking) | Anthropic | 45.0 | 1/1 | LiveCodeBench |
#62 | poolside/Laguna-S-2.1 | poolside | 44.6 | 1/1 | DeepSWE, SWE-bench Pro |
#63 | deepseek-ai/DeepSeek-V4-Flash-0731 | deepseek-ai | 44.4 | 1/1 | DeepSWE |
#64 | inclusionAI/Ring-2.6-1T (high reasoning effort) | inclusionai | 44.0 | 1/1 | SWE-bench Verified |
#65 | ornith-ai/Ornith-1.0-35B (Terminus-2 harness (Harbor framework), parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, avg of 5 runs.) | ornith-ai | 44.0 | 1/1 | Terminal-Bench 2.1 |
#66 | moonshotai/Kimi-K2.5 | moonshotai | 42.9 | 1/1 | SWE-bench Multilingual, SWE-bench Pro |
#67 | stepfun-ai/Step-3.7-Flash | stepfun-ai | 42.6 | 1/1 | SWE-bench Pro, Terminal-Bench 2.1 |
#68 | zai-org/GLM-4.7 | zai-org | 42.0 | 1/1 | SWE-bench Verified |
#69 | MiniMaxAI/MiniMax-M2.5 | minimaxai | 40.6 | 1/1 | SWE-bench Pro |
#70 | Claude-Sonnet-4 (Thinking) | Anthropic | 40.0 | 1/1 | LiveCodeBench |
#71 | thinkingmachines/Inkling (Labeled "Best Harness" on the card (no specific agent/scaffold named). Reported at effort=0.99.) | thinkingmachines | 40.0 | 1/1 | Terminal-Bench 2.1 |
#72 | deepreinforce-ai/Ornith-1.0-35B | deepreinforce-ai | 39.9 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#73 | ornith-ai/Ornith-1.0-35B | ornith-ai | 39.1 | 1/1 | SWE-bench Multilingual |
#74 | ornith-ai/Ornith-1.0-35B (Claude Code 2.1.126 harness, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.) | ornith-ai | 36.0 | 1/1 | Terminal-Bench 2.1 |
#75 | Claude-Opus-4 | Anthropic | 35.0 | 1/1 | LiveCodeBench |
#76 | inclusionAI/Ling-2.6-1T | inclusionai | 34.0 | 1/1 | SWE-bench Verified |
#77 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | nvidia | 32.4 | 1/1 | SWE-bench Multilingual, SWE-bench Verified |
#78 | Claude-Sonnet-4 | Anthropic | 30.0 | 1/1 | LiveCodeBench |
#79 | Qwen/Qwen3.6-35B-A3B | qwen | 29.7 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#80 | moonshotai/Kimi-K2-Thinking | moonshotai | 28.0 | 1/1 | SWE-bench Verified |
#81 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (BF16 checkpoint. Evaluated via NVIDIA's NeMo Evaluator SDK using the Harbor framework, with extended sandboxing via AWS ECS.) | nvidia | 28.0 | 1/1 | Terminal-Bench 2.1 |
#82 | Qwen/Qwen3.6-35B-A3B (Harbor/Terminus-2 harness; 3h timeout, 32 CPU/48 GB RAM; temp=1.0, top_p=0.95, top_k=20, max_tokens=80K, 256K ctx; avg of 5 runs.) | qwen | 26.3 | 1/1 | Terminal-Bench 2.0 |
#83 | poolside/Laguna-M.1 | poolside | 26.1 | 1/1 | SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#84 | DeepSeek-V3 | DeepSeek | 25.0 | 1/1 | LiveCodeBench |
#85 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (NVFP4-quantized checkpoint (same base model as the BF16 release, which scores 56.4). Evaluated via NVIDIA's NeMo Evaluator SDK using the Harbor framework, with extended sandboxing via AWS ECS.) | nvidia | 24.0 | 1/1 | Terminal-Bench 2.1 |
#86 | Qwen/Qwen3.5-122B-A10B | qwen | 23.9 | 1/1 | SWE-bench Verified, Terminal-Bench 2.0 |
#87 | Qwen/Qwen3.8-27B | qwen | 22.2 | 1/1 | DeepSWE |
#88 | Claude-3.5-Sonnet-20241022 | Anthropic | 20.0 | 1/1 | LiveCodeBench |
#89 | meta-models/Muse-Glimmer-30B (Terminus-2 harness, High Reasoning mode.) | meta-models | 20.0 | 1/1 | Terminal-Bench 2.1 |
#90 | Qwen/Qwen3.5-27B | qwen | 18.0 | 1/1 | SWE-bench Verified, Terminal-Bench 2.0 |
#91 | upstage/Solar-Open2-250B | upstage | 18.0 | 1/1 | SWE-bench Verified |
#92 | ornith-ai/Ornith-1.0-9B | ornith-ai | 17.4 | 1/1 | SWE-bench Multilingual |
#93 | ornith-ai/Ornith-1.5-9B | ornith-ai | 17.0 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified |
#94 | ornith-ai/Ornith-1.5-9B (Claude Code 2.1.126 harness, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.) | ornith-ai | 16.0 | 1/1 | Terminal-Bench 2.1 |
#95 | GPT-4O-2024-08-06 | OpenAI | 15.0 | 1/1 | LiveCodeBench |
#96 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 | nvidia | 14.0 | 1/1 | SWE-bench Verified |
#97 | moonshotai/Kimi-K2-Instruct | moonshotai | 13.0 | 1/1 | SWE-bench Multilingual |
#98 | MiniMaxAI/MiniMax-M2 | minimaxai | 12.0 | 1/1 | SWE-bench Verified |
#99 | ornith-ai/Ornith-1.5-9B (Terminus-2 harness (Harbor framework), parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, avg of 5 runs.) | ornith-ai | 12.0 | 1/1 | Terminal-Bench 2.1 |
#100 | poolside/Laguna-XS-2.1 | poolside | 11.1 | 1/1 | SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#101 | papers/2603.00729 (SWE-Agent as harness) | papers | 11.0 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#102 | GPT-4-Turbo-2024-04-09 | OpenAI | 10.0 | 1/1 | LiveCodeBench |
#103 | ornith-ai/Ornith-1.0-9B (Terminus-2 harness (Harbor framework), parser=json, temperature=1.0, top_p=1.0, 128K context, 4h timeout, avg of 5 runs.) | ornith-ai | 8.0 | 1/1 | Terminal-Bench 2.1 |
#104 | poolside/Laguna-XS.2 | poolside | 5.5 | 1/1 | SWE-bench Multilingual, SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#105 | GPT-4O-mini-2024-07-18 | OpenAI | 5.0 | 1/1 | LiveCodeBench |
#106 | internlm/Intern-S2-Preview | internlm | 4.0 | 1/1 | SWE-bench Verified |
#107 | ornith-ai/Ornith-1.0-9B (Claude Code 2.1.126 harness, parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072, avg of 5 runs.) | ornith-ai | 4.0 | 1/1 | Terminal-Bench 2.1 |
#108 | deepreinforce-ai/Ornith-1.0-9B | deepreinforce-ai | 3.4 | 1/1 | SWE-bench Pro, SWE-bench Verified |
#109 | openpangu/openPangu-2.0-Flash (Avg@3, Thinking) | openpangu | 2.0 | 1/1 | SWE-bench Verified |
#110 | Claude-3-Haiku | Anthropic | 0.0 | 1/1 | LiveCodeBench |
#111 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (Evaluated with NVIDIA's NeMo Evaluator SDK (consistent in-house harness used across NVIDIA's release suite).) | nvidia | 0.0 | 1/1 | Terminal-Bench 2.1 |
#112 | inclusionAI/Ling-2.6-flash | inclusionai | -2.0 | 1/1 | SWE-bench Verified |
#113 | CohereLabs/North-Mini-Code-1.0 | coherelabs | -5.3 | 1/1 | SWE-bench Pro, SWE-bench Verified, Terminal-Bench 2.0 |
#114 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (OpenHands harness) | nvidia | -6.0 | 1/1 | SWE-bench Verified |
#115 | Qwen/Qwen3-Coder-Next (agent: Terminus 2) | qwen | -10.5 | 1/1 | Terminal-Bench 2.0 |
#116 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | nvidia | -14.1 | 1/1 | SWE-bench Multilingual, Terminal-Bench 2.0 |
#117 | zai-org/GLM-4.7-Flash | zai-org | -16.0 | 1/1 | SWE-bench Verified |
#118 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (OpenCode harness) | nvidia | -18.0 | 1/1 | SWE-bench Verified |
#119 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | nvidia | -21.8 | 1/1 | SWE-bench Multilingual, SWE-bench Verified |
#120 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nvidia | -22.0 | 1/1 | SWE-bench Multilingual, SWE-bench Verified |
#121 | openpangu/openPangu-2.0-Flash (Avg@3, Non-Thinking) | openpangu | -28.0 | 1/1 | SWE-bench Verified |
#122 | papers/2510.02387 (averaged over 4 runs) | papers | -30.0 | 1/1 | SWE-bench Verified |
#123 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (Codex harness) | nvidia | -32.0 | 1/1 | SWE-bench Verified |