AI directory search

Search across educators, skills, and resources.

Use this when you know the topic you need: Claude Code, MCP, evals, RAG, agents, product, coding, prompting, foundations, or model internals.

8 matches for "benchmarks"

Providers and platforms

OpenRouter profile photo

OpenRouter

OpenRouter docs · Beginner to intermediate

Useful for learning model comparison, latest-family aliases, routing, fallback behavior, agent construction, project-specific evals, and API-compatible experimentation across proprietary and open model families.

Topics

Model routing, Model comparison, Auto Router, Agent SDK, Coding-agent harnesses, GPT models, Claude models, Gemini, Llama, Mistral, DeepSeek, Qwen, API examples, Evaluation

Resources

Designing the hf CLI for agents

Agent tooling and evaluation guide · Hugging Face · Intermediate

You want a measured guide to agent-friendly CLI design, including structured output, retry-safe commands, independent grading, and benchmarks showing where higher-level tools reduce calls and token use.

hugging face, coding agents, hf cli, evals, codex

Simon Willison on GPT-6 Astra

Model analysis · Simon Willison · Intermediate to advanced

You want a concise independent read of GPT-6 Astra's launch claims, benchmark caveats, long-context results, pricing, and early comparison with Claude Fable 5.1 and GPT-5.6 Sol.

gpt-6 astra, model selection, benchmarks, coding agents, long context

How to Choose the Best AI Model (Live, in Your Editor)

Model selection guide · OpenRouter · Intermediate

You want a practical six-step workflow for shortlisting models from live data, testing them on your own prompts, measuring cost per completed task, and choosing or routing from inside a coding assistant.

openrouter, model selection, evals, benchmarks, cost

Perplexity Search Evals

GitHub repo · Perplexity · Intermediate

You want Perplexity's open evaluation suites for benchmarking grounded search quality and comparing search API behavior against other retrieval stacks.

perplexity, search evals, benchmarks, grounded answers, evaluation

OpenRouter benchmarks API

API reference · OpenRouter · Intermediate

You want machine-readable benchmark data for coding, intelligence, or agentic tasks before choosing or routing across model families.

openrouter, benchmarks, model selection, coding, agentic tasks

OpenRouter rankings data API

Data API · OpenRouter · Intermediate

You want current usage and rankings data that reflects what developers are actually using, not only benchmark scores or launch-day marketing.

openrouter, rankings, usage data, benchmarks, model selection

OpenRouter Rankings

Live model rankings · OpenRouter · Beginner to intermediate

You want a current usage-and-benchmark signal for routed models before deciding which providers and families deserve a real evaluation run.

openrouter, rankings, model selection, market signals, benchmarks