Feature 4.1 · ModelMesh AI integration

ModelMesh AI integration

The same 7 capability sliders from the simulator, now applied to real models as they behave in ModelMesh AI. Every bar uses the shared 0–10 visual language so the comparison transfers directly — and every model still reads as AI, which is the honest result at portfolio scale.

Data source
periodic snapshotsnapshot from https://modelmesh-ai.vercel.app as of 2026-09-13. Live vs. snapshot is an explicit decision (PRD §11): snapshot avoids rate limits and latency for this portfolio piece.
Mapping methodology
Heuristic mapping onto the 7 simulator dimensions: reasoning from eval/prompt difficulty, generalization from cross-domain transfer seen in Compare, tool_use from agent/tool success, autonomy from multi-step completion without re-prompting, transfer/learning_speed/self_improvement from architecture limits (no current model redesigns itself). Scores are 0–10 and use the same gates as the simulator — average ≥ 6 + generalization/reasoning ≥ 6 for AGI-like, plus autonomy/self_improvement ≥ 8 and average ≥ 8.5 for ASI-like. No current snapshot scores meet those gates, which is the correct, verifiable result.
Switch to live
To switch this page to live data, expose a CORS-enabled GET /api/public/models on ModelMesh returning [{id,label,provider,scores}] and set NEXT_PUBLIC_MODELMESH_URL to that origin.

GPT-4o · OpenAI

Multimodal chat + tools (ModelMesh Compare/Agent). Strong reasoning, still prompt-bound.

AI · 4.1/10
Generalization
5
Reasoning
7
Autonomy
3
Learning speed
3
Tool use
6
Knowledge transfer
4
Self-improvement
1

AI conceptual profileMatches today's AI: capable within trained areas, but narrow, supervised, or non-transferring somewhere. Raise Generalization and Reasoning to ≥ 6 with an average ≥ 6 to resemble the AGI concept.

Claude 3.5 Sonnet · Anthropic

Long-context reasoning and careful tool use; supervised multi-step agent.

AI · 4.0/10
Generalization
5
Reasoning
7
Autonomy
3
Learning speed
3
Tool use
5
Knowledge transfer
4
Self-improvement
1

AI conceptual profileMatches today's AI: capable within trained areas, but narrow, supervised, or non-transferring somewhere. Raise Generalization and Reasoning to ≥ 6 with an average ≥ 6 to resemble the AGI concept.

Gemini 1.5 Pro · Google

Large-context multimodal; joint text/image/video as seen in ModelMesh.

AI · 4.0/10
Generalization
5
Reasoning
6
Autonomy
3
Learning speed
3
Tool use
5
Knowledge transfer
5
Self-improvement
1

AI conceptual profileMatches today's AI: capable within trained areas, but narrow, supervised, or non-transferring somewhere. Raise Generalization and Reasoning to ≥ 6 with an average ≥ 6 to resemble the AGI concept.

Llama 3.1 405B · Meta

Open-weight dense model; capable chat, limited agent autonomy without scaffolding.

AI · 3.1/10
Generalization
4
Reasoning
6
Autonomy
2
Learning speed
2
Tool use
4
Knowledge transfer
4
Self-improvement
0

AI conceptual profileMatches today's AI: capable within trained areas, but narrow, supervised, or non-transferring somewhere. Raise Generalization and Reasoning to ≥ 6 with an average ≥ 6 to resemble the AGI concept.

Mistral Large 2 · Mistral AI

Efficient frontier chat; tool use via ModelMesh agent, not self-directed.

AI · 3.0/10
Generalization
4
Reasoning
6
Autonomy
2
Learning speed
2
Tool use
4
Knowledge transfer
3
Self-improvement
0

AI conceptual profileMatches today's AI: capable within trained areas, but narrow, supervised, or non-transferring somewhere. Raise Generalization and Reasoning to ≥ 6 with an average ≥ 6 to resemble the AGI concept.

Conceptual model — these scores are a teaching mapping onto the simulator's 7 dimensions, not benchmark citations. All 5 models read as AI under the same gates the simulator uses (average ≥ 6 + generalization/reasoning ≥ 6 for AGI-like). That is the portfolio narrative working as intended: ModelMesh shows what current models can do; IntelliGenesis shows what those capabilities would need to become to resemble AGI/ASI.