- 作者
- DoDoT Data Lab
- Provider
- BazaarLink
- 交付方式
- 站内运行
每 10 分钟本地快照
模型智力监控
展示 DoDoT 定时保存的 BazaarLink 模型检测快照;手动刷新会重新采集并真实入库。这是来源方检测结果,不是 DoDoT 自建评测分数。
- 基准模型
- 30
- 总测试数
- 2.6万
- 24H 测试
- 866
- 24H 合格率
- 27.4%
最近 24 小时
仅展示最新合格率最高的 8 个模型合格率趋势
基于最近 142 个真实采集快照
GPT-5.6 LunaDeepSeek V4 FlashGPT-5.6 SolClaude Fable 5GLM-5.2Grok 4.5DeepSeek V4 ProGemini 3.6 Flash
模型总测试数24H 测试(合格率)站点最后更新
GPT-5.6 Solopenai/gpt-5.6-sol
6,855
428(35.3%)
1,738
GPT-5.6 Terraopenai/gpt-5.6-terra
816
40(17.5%)
414
anthropic/claude-opus-4.6
2,154
12(16.7%)
802
GPT-5.6 Lunaopenai/gpt-5.6-luna
374
32(43.8%)
242
Claude Opus 5anthropic/claude-opus-5
1,251
103(19.4%)
438
Claude Fable 5anthropic/claude-fable-5
1,473
53(32.1%)
601
anthropic/claude-opus-4.8
3,422
96(5.2%)
1,043
GPT-5.5openai/gpt-5.5
5,212
17(5.9%)
1,740
DeepSeek V4 Prodeepseek/deepseek-v4-pro
275
12(25.0%)
129
DeepSeek V4 Flashdeepseek/deepseek-v4-flash
332
22(36.4%)
164
GLM-5.2z-ai/glm-5.2
495
13(30.8%)
199
anthropic/claude-sonnet-4.6
456
1(0.0%)
258
Claude Sonnet 5anthropic/claude-sonnet-5
297
10(10.0%)
164
anthropic/claude-opus-4.7
1,340
1(0.0%)
600
Grok 4.5x-ai/grok-4.5
291
10(30.0%)
157
GPT-5.4openai/gpt-5.4
785
4(0.0%)
466
Gemini 3.6 Flashgoogle/gemini-3.6-flash
23
4(25.0%)
16
GLM-5.1z-ai/glm-5.1
137
0(尚无判定)
77
Grok 4.3x-ai/grok-4.3
12
0(尚无判定)
9
GLM-5z-ai/glm-5
25
0(尚无判定)
18
openai/gpt-4o
26
0(尚无判定)
13
anthropic/claude-haiku-4.5
140
0(尚无判定)
74
GPT-5.4 Miniopenai/gpt-5.4-mini
95
0(尚无判定)
66
Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite
5
0(尚无判定)
3
DeepSeek V3.2deepseek/deepseek-v3.2
4
0(尚无判定)
4
GPT-4openai/gpt-4
13
0(尚无判定)
10
GPT-5.2openai/gpt-5.2
37
0(尚无判定)
24
GPT-5.3 Codexopenai/gpt-5.3-codex
65
0(尚无判定)
47
Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview
1
8(0.0%)
1
GPT-3.5 Turboopenai/gpt-3.5-turbo
1
0(尚无判定)
1