Сводный рейтинг Neira объединяет независимые тесты, сохраняя точные версии моделей и покрытие категорий.
Последний полный снимок: 23 авг. 2026 г., 06:02. Чат 30%, знания и рассуждение 20%, код 40%, вызов инструментов 10%.
Оценка 0–100 нормализована по позиции в каждом источнике, поэтому исходные шкалы не смешиваются.
| Ранг | Модель | Разработчик | Оценка | Покрытие | Источники |
|---|---|---|---|---|---|
#1 | claude-fable-5-max-effort | Anthropic | 100.0 | 1/1 | LiveBench |
#2 | ornith-ai/Ornith-1.5-397B (With tools.) | ornith-ai | 100.0 | 1/1 | Humanity's Last Exam |
Открытые веса · 4495 моделей
#3 |
| moonshotai/Kimi-K3 |
| moonshotai |
| 99.2 |
| 1/1 |
| GPQA, Humanity's Last Exam |
#4 | ornith-ai/Ornith-1.5-397B | ornith-ai | 98.3 | 1/1 | GPQA |
#5 | gpt-5.6-sol-max | OpenAI | 97.4 | 1/1 | LiveBench |
#6 | zai-org/GLM-5.2 (With tools) | zai-org | 96.7 | 1/1 | Humanity's Last Exam |
#7 | internlm/Intern-S2-Preview | internlm | 96.3 | 1/1 | MMLU-Pro |
#8 | moonshotai/Kimi-K2.6 (With tools) | moonshotai | 95.0 | 1/1 | Humanity's Last Exam |
#9 | gpt-5.5-xhigh | OpenAI | 94.9 | 1/1 | LiveBench |
#10 | FINAL-Bench/Darwin-398B-JGOS (greedy decoding (temperature=0), single-sample (no voting / no test-time engine), max_tokens=16384, options shuffled seed=42; hardware: NVIDIA B200 x6 (TP2 x PP3), vLLM bfloat16) | final-bench | 93.3 | 1/1 | GPQA |
#11 | gemini-3.1-pro-preview-high | 92.3 | 1/1 | LiveBench |
#12 | dots-studio/dots3-note-prev (Reported as 'HLE w/ tool' — evaluated with tool/browsing access enabled, not the closed-book default.) | dots-studio | 90.0 | 1/1 | Humanity's Last Exam |
#13 | tencent/Hy3 | tencent | 90.0 | 1/1 | GPQA, Humanity's Last Exam |
#14 | gemini-3.7-flash-high | 89.7 | 1/1 | LiveBench |
#15 | zai-org/GLM-5.1 (With tools) | zai-org | 88.3 | 1/1 | Humanity's Last Exam |
#16 | gpt-5.4-xhigh | OpenAI | 87.2 | 1/1 | LiveBench |
#17 | zai-org/GLM-5 (With tools) | zai-org | 86.7 | 1/1 | Humanity's Last Exam |
#18 | moonshotai/Kimi-K2.5 (With tools) | moonshotai | 85.0 | 1/1 | Humanity's Last Exam |
#19 | claude-opus-5-max-effort | Anthropic | 84.6 | 1/1 | LiveBench |
#20 | FINAL-Bench/Darwin-28B-REASON (Darwin-DELPHI test-time engine, Pass@1) | final-bench | 83.3 | 1/1 | GPQA |
#21 | Qwen/Qwen3.5-27B (with tools) | qwen | 83.3 | 1/1 | Humanity's Last Exam |
#22 | Qwen/Qwen3.8-2.4T-A95B | qwen | 82.5 | 1/1 | GPQA, Humanity's Last Exam |
#23 | grok-4.6 | xAI | 82.1 | 1/1 | LiveBench |
#24 | Qwen/Qwen3.5-397B-A17B (With tools) | qwen | 81.7 | 1/1 | Humanity's Last Exam |
#25 | Qwen/Qwen3.8-27B | qwen | 81.7 | 1/1 | GPQA |
#26 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | nvidia | 81.5 | 1/1 | MMLU-Pro |
#27 | ornith-ai/Ornith-1.5-35B-A3B | ornith-ai | 80.0 | 1/1 | GPQA |
#28 | stepfun-ai/Step-3.7-Flash (With tools) | stepfun-ai | 80.0 | 1/1 | Humanity's Last Exam |
#29 | zai-org/GLM-5.2 | zai-org | 80.0 | 1/1 | GPQA, Humanity's Last Exam |
#30 | qwen3.8-max | Alibaba | 79.5 | 1/1 | LiveBench |
#31 | XiaomiMiMo/MiMo-V2.5-Pro (with tools) | xiaomimimo | 78.3 | 1/1 | Humanity's Last Exam |
#32 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 | nvidia | 77.8 | 1/1 | MMLU-Pro |
#33 | thinkingmachines/Inkling | thinkingmachines | 77.8 | 1/1 | GPQA, Humanity's Last Exam, MathArena AIME 2026 |
#34 | gpt-5.6-terra-max | OpenAI | 76.9 | 1/1 | LiveBench |
#35 | FINAL-Bench/Darwin-36B-Opus (Standard inference, Pass@1) | final-bench | 76.7 | 1/1 | GPQA |
#36 | InternScience/Agents-A1 (With tools) | internscience | 76.7 | 1/1 | Humanity's Last Exam |
#37 | deepseek-ai/DeepSeek-V4-Pro | deepseek-ai | 76.4 | 1/1 | GPQA, GSM8K, Humanity's Last Exam, MMLU-Pro |
#38 | moonshotai/Kimi-K2.5 | moonshotai | 75.1 | 1/1 | GPQA, MMLU-Pro |
#39 | moonshotai/Kimi-K2.6 | moonshotai | 75.1 | 1/1 | GPQA, Humanity's Last Exam, MathArena AIME 2026 |
#40 | FINAL-Bench/Darwin-60B-DUO (DELPHI cascade system, Pass@1) | final-bench | 75.0 | 1/1 | GPQA |
#41 | Qwen/Qwen3.5-122B-A10B (with tools) | qwen | 75.0 | 1/1 | Humanity's Last Exam |
#42 | kimi-k3 | moonshot | 74.4 | 1/1 | LiveBench |
#43 | inclusionAI/Ring-2.6-1T (xhigh reasoning effort, Pass@1) | inclusionai | 73.3 | 1/1 | GPQA |
#44 | deepseek-v4-pro-0813 | DeepSeek | 71.8 | 1/1 | LiveBench |
#45 | moonshotai/Kimi-K2-Thinking (With tools) | moonshotai | 71.7 | 1/1 | Humanity's Last Exam |
#46 | ornith-ai/Ornith-1.5-397B (No tools.) | ornith-ai | 70.0 | 1/1 | Humanity's Last Exam |
#47 | Qwen/Qwen3.5-397B-A17B | qwen | 69.8 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro |
#48 | grok-4.5 | xAI | 69.2 | 1/1 | LiveBench |
#49 | deepseek-ai/DeepSeek-V4-Flash | deepseek-ai | 66.8 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro |
#50 | claude-opus-4-8-max-effort | Anthropic | 66.7 | 1/1 | LiveBench |
#51 | zai-org/GLM-4.7 (With tools) | zai-org | 66.7 | 1/1 | Humanity's Last Exam |
#52 | deepseek-v4-flash-vision-exp | DeepSeek | 64.1 | 1/1 | LiveBench |
#53 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (With tools) | nvidia | 61.7 | 1/1 | Humanity's Last Exam |
#54 | tencent/Hy3-preview | tencent | 61.7 | 1/1 | GPQA |
#55 | claude-opus-4-7-xhigh-effort | Anthropic | 61.5 | 1/1 | LiveBench |
#56 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (With tools) | nvidia | 60.0 | 1/1 | Humanity's Last Exam |
#57 | gemini-3.5-flash-high | 59.0 | 1/1 | LiveBench |
#58 | thinkingmachines/Inkling-Small | thinkingmachines | 57.9 | 1/1 | GPQA, MathArena AIME 2026 |
#59 | FINAL-Bench/Darwin-27B-Opus (Standard inference, Pass@1) | final-bench | 56.7 | 1/1 | GPQA |
#60 | deepseek-v4-flash-0731 | DeepSeek | 56.4 | 1/1 | LiveBench |
#61 | Qwen/Qwen3.5-122B-A10B | qwen | 55.0 | 1/1 | GPQA |
#62 | gemini-3.6-flash-high | 53.8 | 1/1 | LiveBench |
#63 | ornith-ai/Ornith-1.5-9B | ornith-ai | 53.3 | 1/1 | GPQA |
#64 | PolarSeeker/OpenSeeker-v2-30B-SFT | polarseeker | 53.3 | 1/1 | Humanity's Last Exam |
#65 | FINAL-Bench/Ourbox-35B-JGOS (majority-of-8+ (maj@8 with 16-vote tiebreaker)) | final-bench | 51.7 | 1/1 | GPQA |
#66 | ornith-ai/Ornith-1.5-35B-A3B (With tools.) | ornith-ai | 51.7 | 1/1 | Humanity's Last Exam |
#67 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4 (No tools) | nvidia | 51.7 | 1/1 | GPQA, Humanity's Last Exam |
#68 | gpt-5.2-2025-12-11-high | OpenAI | 51.3 | 1/1 | LiveBench |
#69 | microsoft/Phi-3-medium-4k-instruct (8-shot CoT) | microsoft | 50.0 | 1/1 | GSM8K |
#70 | thinkingmachines/Inkling-Small (text only) | thinkingmachines | 50.0 | 1/1 | Humanity's Last Exam |
#71 | qwen3.7-max | Alibaba | 48.7 | 1/1 | LiveBench |
#72 | deepseek-ai/DeepSeek-R1-0528 | deepseek-ai | 48.1 | 1/1 | MMLU-Pro |
#73 | upstage/Solar-Open2-250B | upstage | 47.9 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro, MathArena AIME 2026 |
#74 | nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (No tools) | nvidia | 46.7 | 1/1 | GPQA, Humanity's Last Exam |
#75 | Qwen/Qwen3.8-27B (Judged by GPT-4o.) | qwen | 46.7 | 1/1 | Humanity's Last Exam |
#76 | claude-opus-4-6-thinking-auto-high-effort | Anthropic | 46.2 | 1/1 | LiveBench |
#77 | inclusionAI/Ring-2.6-1T (xhigh reasoning effort, Mean@64) | inclusionai | 46.2 | 1/1 | MathArena AIME 2026 |
#78 | zai-org/GLM-5 | zai-org | 45.8 | 1/1 | GPQA, Humanity's Last Exam |
#79 | claude-sonnet-5-xhigh-effort | Anthropic | 43.6 | 1/1 | LiveBench |
#80 | FINAL-Bench/Darwin-31B-Opus (Standard inference, Pass@1) | final-bench | 43.3 | 1/1 | GPQA |
#81 | ornith-ai/Ornith-1.5-9B (With tools.) | ornith-ai | 43.3 | 1/1 | Humanity's Last Exam |
#82 | Qwen/Qwen2-72B | qwen | 41.7 | 1/1 | GSM8K |
#83 | tencent/Hy3-preview (Text-only) | tencent | 41.7 | 1/1 | Humanity's Last Exam |
#84 | Qwen/Qwen3.6-27B | qwen | 41.1 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro, MathArena AIME 2026 |
#85 | qwen3.8-27b | Alibaba | 41.0 | 1/1 | LiveBench |
#86 | zai-org/GLM-5.1 | zai-org | 39.9 | 1/1 | GPQA, Humanity's Last Exam, MathArena AIME 2026 |
#87 | claude-sonnet-4-6-thinking-auto-medium-effort | Anthropic | 38.5 | 1/1 | LiveBench |
#88 | Qwen/Qwen3.5-27B | qwen | 38.3 | 1/1 | GPQA |
#89 | deepseek-v4-pro | DeepSeek | 35.9 | 1/1 | LiveBench |
#90 | zai-org/GLM-4.7 | zai-org | 35.0 | 1/1 | GPQA, Humanity's Last Exam |
#91 | claude-opus-4-5-20251101-thinking-64k-high-effort | Anthropic | 33.3 | 1/1 | LiveBench |
#92 | FINAL-Bench/Darwin-4B-David (4B class Gen-2 evolution, Pass@1) | final-bench | 33.3 | 1/1 | GPQA |
#93 | ornith-ai/Ornith-1.5-35B-A3B (No tools.) | ornith-ai | 31.7 | 1/1 | Humanity's Last Exam |
#94 | gpt-5.2-codex | OpenAI | 30.8 | 1/1 | LiveBench |
#95 | Qwen/Qwen3.5-122B-A10B (chain of thought) | qwen | 30.0 | 1/1 | Humanity's Last Exam |
#96 | FINAL-Bench/Darwin-9B-NEG (9B class, standard inference, Pass@1) | final-bench | 28.3 | 1/1 | GPQA |
#97 | gpt-5.6-luna-max | OpenAI | 28.2 | 1/1 | LiveBench |
#98 | JGOS-Model/JGOS-31B-Citizen (maj@8 + DELPHI + near-miss maj@32-64 weighted vote, Pass@1 (167/198)) | jgos-model | 26.7 | 1/1 | GPQA |
#99 | moonshotai/Kimi-K2-Thinking | moonshotai | 25.8 | 1/1 | GPQA, Humanity's Last Exam |
#100 | glm-5.2 | z-ai | 25.6 | 1/1 | LiveBench |
#101 | openpangu/openPangu-2.0-Flash (Avg@4, Thinking) | openpangu | 25.0 | 1/1 | GPQA |
#102 | Qwen/Qwen3.5-27B (chain of thought) | qwen | 25.0 | 1/1 | Humanity's Last Exam |
#103 | gpt-5.4-nano-xhigh | OpenAI | 23.1 | 1/1 | LiveBench |
#104 | kimi-k2.6-thinking | moonshot | 20.5 | 1/1 | LiveBench |
#105 | Qwen/Qwen3.6-35B-A3B | qwen | 18.9 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro, MathArena AIME 2026 |
#106 | grok-build-0.1 | xAI | 17.9 | 1/1 | LiveBench |
#107 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (With tools) | nvidia | 17.5 | 1/1 | GPQA, Humanity's Last Exam |
#108 | MiniMaxAI/MiniMax-M2.5 | minimaxai | 16.7 | 1/1 | GPQA, Humanity's Last Exam |
#109 | qwen3.6-plus | Alibaba | 15.4 | 1/1 | LiveBench |
#110 | meta-models/Muse-Glimmer-30B | meta-models | 14.0 | 1/1 | GPQA, Humanity's Last Exam, MathArena AIME 2026 |
#111 | kimi-k2.7-code | moonshot | 12.8 | 1/1 | LiveBench |
#112 | deepseek-v4-flash | DeepSeek | 10.3 | 1/1 | LiveBench |
#113 | internlm/internlm2_5-7b-chat (0-shot CoT) | internlm | 8.3 | 1/1 | GSM8K |
#114 | MiniMaxAI/MiniMax-M2.1 | minimaxai | 8.3 | 1/1 | Humanity's Last Exam |
#115 | gpt-5.4-mini-xhigh | OpenAI | 7.7 | 1/1 | LiveBench |
#116 | XiaomiMiMo/MiMo-V2-Flash | xiaomimimo | 6.7 | 1/1 | Humanity's Last Exam |
#117 | nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | nvidia | 6.0 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro |
#118 | grok-4.3 | xAI | 5.1 | 1/1 | LiveBench |
#119 | internlm/Intern-S2-Preview (Text only) | internlm | 3.3 | 1/1 | Humanity's Last Exam |
#120 | qwen3.6-27b | Alibaba | 2.6 | 1/1 | LiveBench |
#121 | gemini-3.5-flash-lite-high | 0.0 | 1/1 | LiveBench |
#122 | LGAI-EXAONE/K-EXAONE-236B-A23B | lgai-exaone | 0.0 | 1/1 | GPQA, Humanity's Last Exam, MMLU-Pro |
#123 | meituan-longcat/LongCat-Flash-Thinking-2601 | meituan-longcat | 0.0 | 1/1 | GPQA |
#124 | microsoft/Phi-3-mini-4k-instruct (8-shot CoT) | microsoft | 0.0 | 1/1 | GSM8K |
#125 | ornith-ai/Ornith-1.5-9B (No tools.) | ornith-ai | 0.0 | 1/1 | Humanity's Last Exam |
#126 | deepseek-ai/DeepSeek-R1 | deepseek-ai | -0.6 | 1/1 | GPQA, MMLU-Pro |
#127 | inclusionAI/Ling-3.0-flash | inclusionai | -2.7 | 1/1 | Humanity's Last Exam, MathArena AIME 2026 |
#128 | openpangu/openPangu-2.0-Flash (Avg@4, Non-Thinking) | openpangu | -5.0 | 1/1 | GPQA |
#129 | openpangu/openPangu-2.0-Flash (Thinking) | openpangu | -7.7 | 1/1 | MathArena AIME 2026 |
#130 | LGAI-EXAONE/EXAONE-4.5-33B | lgai-exaone | -8.3 | 1/1 | GPQA, MMLU-Pro, MathArena AIME 2026 |
#131 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | nvidia | -11.1 | 1/1 | MMLU-Pro |
#132 | XiaomiMiMo/MiMo-V2.5-Pro | xiaomimimo | -11.2 | 1/1 | GPQA, GSM8K, MMLU-Pro |
#133 | MiniMaxAI/MiniMax-M2 | minimaxai | -13.7 | 1/1 | Humanity's Last Exam, MMLU-Pro |
#134 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | nvidia | -18.2 | 1/1 | GPQA, MMLU-Pro |
#135 | skt/A.X-K1 (English MMLU-Pro score reported in Thinking Mode.) | skt | -18.5 | 1/1 | MMLU-Pro |
#136 | nvidia/Nemotron-Cascade-2-30B-A3B | nvidia | -20.0 | 1/1 | GPQA |
#137 | zai-org/GLM-4.7-Flash | zai-org | -20.8 | 1/1 | GPQA, Humanity's Last Exam |
#138 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (Text-only, no tools.) | nvidia | -21.7 | 1/1 | Humanity's Last Exam |
#139 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 (No tools.) | nvidia | -23.3 | 1/1 | GPQA |
#140 | Qwen/Qwen2-7B (4-shot) | qwen | -25.0 | 1/1 | GSM8K |
#141 | deepseek-ai/DeepSeek-V3-0324 | deepseek-ai | -25.9 | 1/1 | MMLU-Pro |
#142 | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 (Text-only, no tools — HLE's default includes image questions, so this excludes the multimodal subset.) | nvidia | -26.7 | 1/1 | Humanity's Last Exam |
#143 | HelpingAI/Dhanishtha-2.0-0126 | helpingai | -28.3 | 1/1 | Humanity's Last Exam |
#144 | jdopensource/JoyAI-LLM-Flash | jdopensource | -30.0 | 1/1 | GPQA, MMLU-Pro |
#145 | skt/A.X-K1 (GPQA-Diamond score reported in Thinking Mode.) | skt | -30.0 | 1/1 | GPQA |
#146 | internlm/internlm2-chat-20b | internlm | -33.3 | 1/1 | GSM8K |
#147 | ibm-granite/granite-4.1-30b | ibm-granite | -34.0 | 1/1 | GPQA, GSM8K, MMLU-Pro |
#148 | skt/A.X-K1 (Humanity's Last Exam score reported in Thinking Mode.) | skt | -35.0 | 1/1 | Humanity's Last Exam |
#149 | deepseek-ai/DeepSeek-V3 | deepseek-ai | -40.7 | 1/1 | GSM8K, MMLU-Pro |
#150 | deepseek-ai/DeepSeek-V2 | deepseek-ai | -41.7 | 1/1 | GSM8K |
#151 | mistralai/Mistral-Small-4-119B-2603 (reasoning: high) | mistralai | -41.7 | 1/1 | GPQA |
#152 | Qwen/Qwen3-4B-Thinking-2507 | qwen | -51.7 | 1/1 | GPQA |
#153 | meituan-longcat/LongCat-Flash-Lite | meituan-longcat | -55.6 | 1/1 | MMLU-Pro |
#154 | ibm-granite/granite-4.1-8b | ibm-granite | -58.9 | 1/1 | GPQA, GSM8K, MMLU-Pro |
#155 | openpangu/openPangu-2.0-Flash (Non-Thinking) | openpangu | -61.5 | 1/1 | MathArena AIME 2026 |
#156 | Qwen/Qwen3-4B-Instruct-2507 | qwen | -63.7 | 1/1 | GPQA, MMLU-Pro |
#157 | inclusionAI/Ling-2.6-flash | inclusionai | -76.9 | 1/1 | MathArena AIME 2026 |
#158 | ibm-granite/granite-4.1-3b | ibm-granite | -86.4 | 1/1 | GPQA, GSM8K, MMLU-Pro |