Evidence Explorer

公開している claim(強い主張)を証拠区分・タイプで絞り込み。無根拠の強い主張は公開していません(§8.4)。

フィルター

各 claim 行の # リンクで permalink(例: #claim-gemini-fast-cheap)がコピーできます。

Claims(10件)

permalinkclaimIdtexttypegradestatusscopecheckedAt
# claim-hy3-free-limited Tencent Hy3 には期間限定の無料バリアント (tencent/hy3:free) が存在する。 spec A verified model 2026-07-14
# claim-hy3-free-expires tencent/hy3:free の無料提供は 2026-07-21 に期限切れとなる。 spec A verified model 2026-07-14
# claim-deepseek-coding-strong DeepSeek V3.2 はコード生成・リファクタに向くと編集者は判断する。 editorial E unverified model 2026-07-14
# claim-claude-japanese Claude Sonnet 5 は日本語・長文・慎重な業務に向くと編集者は判断する。 editorial E qualified model 2026-07-14
# claim-gemini-fast-cheap Gemini 3.5 Flash は低価格・高速・長文コンテキスト向けと編集者は判断する。 editorial E qualified model 2026-07-14
# claim-context-not-effective 公式の context window 長は、実効的に使いこなせる長さ(effective context)と同義ではない。 caveat D verified global 2026-07-14
# claim-llm-judge-bias LLM-as-a-judge 単独では性能順位を断定的に公開すべきではない(position 等の bias あり)。 methodology D verified global 2026-07-14
# claim-swebench-coding-complex real-world のコーディング課題は、単純なコード生成より複合的である。 methodology D verified use-case 2026-07-14
# claim-arena-preference-not-correctness human preference(人間の好み)と objective correctness(客観的正解率)は別の評価軸である。 methodology D verified global 2026-07-14
# claim-helm-multimetric モデル評価は単一の accuracy だけでなく、use-case と複数の metrics に分けて行うべきである。 methodology D verified global 2026-07-14