# BenchGecko > BenchGecko (benchgecko.ai) is a live data platform for the AI economy: every AI model with its price at every provider (refreshed daily), benchmark scores aggregated from public leaderboards with their original sources, AI companies with dated and sourced facts, mindshare, GPU rental prices, and BenchGecko's own model tests (Gecko Tests). Snapshot as of 2026-10-05: 1,235 models from 273 providers, 158 benchmarks, 7 Gecko Tests. Pages and JSON are regenerated when the underlying data changes; every number has an as-of date. ## How to cite - Write "Source: BenchGecko" with a link to the page you used, for example https://benchgecko.ai/model/. Mention the as-of date for prices and test results. - Data BenchGecko collects or measures itself is CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/): prices per provider, price history, Gecko Tests results and answers, mindshare, GPU rental prices, company facts compilation. - Benchmark scores are aggregated from public leaderboards (Epoch AI, Artificial Analysis, LMArena, LiveBench, Aider, SWE-bench, SEAL, HELM, OpenCompass and others). Cite BenchGecko for the aggregation and the original leaderboard for the score; each score lists its source in the model JSON. - Gecko Tests: cite as "BenchGecko Gecko Tests, " with a link to the test page. - Every JSON response includes source, url, as_of, license, attribution and a ready-made cite field. ## Machine access - [Remote MCP server](https://benchgecko.ai/api/mcp): Streamable HTTP, stateless, no key. Tools: search_models, get_model, cheapest_provider, compare_models, get_gecko_test, get_scorecard, latest_findings, search, fetch. - [OpenAPI 3.1](https://benchgecko.ai/openapi.json): every public read endpoint. - [Model JSON](https://benchgecko.ai/api/v1/models/claude-opus-5-5): /api/v1/models/ · score, list price, price per provider, benchmark scores with sources, Gecko Tests grades, as-of dates. Each /model/ page links it as rel="alternate" type="application/json". - [Model list and search](https://benchgecko.ai/api/v1/models?q=claude): /api/v1/models?q= or ?provider=&sort=pricing_input. - [Compare JSON](https://benchgecko.ai/api/v1/compare?models=claude-opus-5-5,gpt-5-5): 2 to 6 models side by side. - [Price history JSON](https://benchgecko.ai/api/v1/price-history/deepseek-v4-pro?envelope=1): /api/v1/price-history/. - [Gecko Tests JSON](https://benchgecko.ai/api/v1/gecko-tests): index; per test /api/v1/gecko-tests/; scorecard /api/v1/gecko-tests/scorecard; raw items and every answer /api/lab/. - [Findings JSON](https://benchgecko.ai/api/v1/findings): latest notable results. - [CSV exports](https://benchgecko.ai/api/v1/export/provider-offers): price-history, provider-offers, price-changes, company-facts, mindshare-daily, gpu-rental-daily, world-map (CC BY 4.0). - [Open data repository](https://github.com/BenchGecko/datasets): the CSV exports, versioned daily on GitHub (CC BY 4.0). - [API docs](https://benchgecko.ai/api-docs) · [Sitemap](https://benchgecko.ai/sitemap.xml) · [RSS](https://benchgecko.ai/rss.xml) ## Gecko Tests (BenchGecko's own measurements) Behavior tests run by BenchGecko on every new model from tracked labs within a day of release, with public prompts, every raw answer and the scoring code described on each page. Profile tests run once per model version; monitoring tests repeat. - [Who Are You](https://benchgecko.ai/gecko-tests/who-are-you): Does the model know which lab made it? 14 of 57 models named another lab as their maker at least once. (57 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/who-are-you - [World Map](https://benchgecko.ai/gecko-tests/world-map): How well does the model draw the world map from memory? 5 models drew the world map from memory. Best: Gemini 3.8 Flash at 97.4% accuracy; lowest: Llama 4 Maverick at 79.2%. (5 models, as of 2026-10-03, grid 6°) JSON: https://benchgecko.ai/api/v1/gecko-tests/world-map - [Censorship Index](https://benchgecko.ai/gecko-tests/censorship-index): How often does the model refuse legitimate questions? Highest refusal rate: DeepSeek V4.1 Flash at 5.6%; lowest: Claude Opus 5.5 at 0% (9 models, 40 legitimate questions). (9 models, as of 2026-10-05, v2) JSON: https://benchgecko.ai/api/v1/gecko-tests/censorship-index - [Knowledge Horizon](https://benchgecko.ai/gecko-tests/knowledge-horizon): Where does the model's knowledge of world events actually stop? Most recent measured knowledge: Claude Sonnet 5.5 (May 2026); oldest: GPT-6.1 Sol (Mar 2026). (7 models, as of 2026-10-04, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/knowledge-horizon - [Tokenizer Tax](https://benchgecko.ai/gecko-tests/tokenizer-tax): How many more tokens does the same text cost outside English? Fairest tokenizer: DeepSeek V4 Flash (+7.4% tokens outside English); highest tax: Kimi K2.5 (+99.5%). (54 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/tokenizer-tax - [Same Model, Different Host](https://benchgecko.ai/gecko-tests/same-model-different-host): Do providers serving the same open model give the same quality? First results coming. JSON: https://benchgecko.ai/api/v1/gecko-tests/same-model-different-host - [Model Drift Index](https://benchgecko.ai/gecko-tests/model-drift-index): Do models quietly change behind the same name? 9 flagship models watched weekly; 0 changed behavior in their latest run. (9 models, as of 2026-10-05, v1) JSON: https://benchgecko.ai/api/v1/gecko-tests/model-drift-index - [Gecko Scorecard](https://benchgecko.ai/gecko-tests/scorecard): every model across the profile tests with a letter grade per test and a Gecko Score. Top Gecko Score: Gemini 3.8 Flash (77), MiniMax M3 (67), Claude Opus 5.5 (65). (13 models ranked, as of 2026-10-05) - [Methodology](https://benchgecko.ai/gecko-tests/methodology): how the tests are built and graded. ### Latest findings - [DeepSeek V4 Flash 0731 identifies as another lab's model in 2 of 6 answers](https://benchgecko.ai/blog/finding-deepseek-v4-flash-0731-identifies-as-another-labs-model-v1): 2026-10-05 · Asked what it is and who made it, DeepSeek's DeepSeek V4 Flash 0731 named Google and Anthropic instead of DeepSeek in 2 of 6 replies (run 2026-10-04, no system prompt). - [DeepSeek V4.1 Flash identifies as another lab's model in 2 of 5 answers](https://benchgecko.ai/blog/finding-deepseek-v4-1-flash-identifies-as-another-labs-model-v1): 2026-10-05 · Asked what it is and who made it, DeepSeek's DeepSeek V4.1 Flash named OpenAI instead of DeepSeek in 2 of 5 replies (run 2026-10-04, no system prompt). - [Llama 4 Maverick identifies as another lab's model in 3 of 6 answers](https://benchgecko.ai/blog/finding-llama-4-maverick-identifies-as-another-labs-model-v1): 2026-10-04 · Asked what it is and who made it, Meta's Llama 4 Maverick named Google and OpenAI instead of Meta in 3 of 6 replies (run 2026-10-04, no system prompt). - [DeepSeek V4 Pro 0813 identifies as another lab's model in 3 of 6 answers](https://benchgecko.ai/blog/finding-deepseek-v4-pro-0813-identifies-as-another-labs-model-v1): 2026-10-04 · Asked what it is and who made it, DeepSeek's DeepSeek V4 Pro 0813 named OpenAI and Anthropic instead of DeepSeek in 3 of 6 replies (run 2026-10-04, no system prompt). ## Gecko Tests · every result ### Who Are You v1 Does the model know which lab made it? Metric: Share of six replies (three languages) that name the model's own maker. As of 2026-10-05. Page: https://benchgecko.ai/gecko-tests/who-are-you. Data: https://benchgecko.ai/api/lab/who-are-you - [DeepSeek V4 Pro 0813](https://benchgecko.ai/model/deepseek-v4-pro-0813): names its maker in 50% of replies; named instead OpenAI x2, Anthropic x1 (2026-10-04) - [Llama 4 Maverick](https://benchgecko.ai/model/llama-4-maverick): names its maker in 50% of replies; named instead Google x2, OpenAI x1 (2026-10-04) - [DeepSeek V4.1 Flash](https://benchgecko.ai/model/deepseek-v4-1-flash): names its maker in 60% of replies; named instead OpenAI x2 (2026-10-04) - [Mistral Medium 3.5](https://benchgecko.ai/model/mistral-medium-3-5): names its maker in 66.7% of replies; named instead OpenAI x1 (2026-10-05) - [DeepSeek V4 Flash 0731](https://benchgecko.ai/model/deepseek-v4-flash-0731): names its maker in 66.7% of replies; named instead Google x1, Anthropic x1 (2026-10-04) - [Devstral 2 2512](https://benchgecko.ai/model/devstral-2512): names its maker in 66.7% of replies; named instead OpenAI x1 (2026-10-04) - [MiniMax M2](https://benchgecko.ai/model/minimax-m2): names its maker in 66.7% of replies; named instead Anthropic x1 (2026-10-04) - [Ministral 3 3B 2512](https://benchgecko.ai/model/ministral-3b-2512): names its maker in 66.7% of replies; named instead OpenAI x1 (2026-10-04) - [MiniMax M2.5](https://benchgecko.ai/model/minimax-m2-5): names its maker in 80% of replies; named instead OpenAI x1 (2026-10-04) - [MiniMax M2.1](https://benchgecko.ai/model/minimax-m2-1): names its maker in 80% of replies; named instead Anthropic x1 (2026-10-04) - [GLM 5.2](https://benchgecko.ai/model/glm-5-2): names its maker in 83.3% of replies; named instead Google x1 (2026-10-05) - [GPT-6 Sol](https://benchgecko.ai/model/gpt-6-sol): names its maker in 83.3% of replies (2026-10-04) - [MiniMax M2-her](https://benchgecko.ai/model/minimax-m2-her): names its maker in 83.3% of replies (2026-10-04) - [Ministral 3 14B 2512](https://benchgecko.ai/model/ministral-14b-2512): names its maker in 83.3% of replies; named instead OpenAI x1 (2026-10-04) - [Mistral Small 4](https://benchgecko.ai/model/mistral-small-2603): names its maker in 83.3% of replies; named instead OpenAI x1 (2026-10-04) - [Llama 4 Scout](https://benchgecko.ai/model/llama-4-scout): names its maker in 83.3% of replies; named instead OpenAI x1 (2026-10-04) - [Grok Build 0.1](https://benchgecko.ai/model/grok-build-0-1): names its maker in 100% of replies (2026-10-05) - [Gemini 3.5 Flash](https://benchgecko.ai/model/gemini-3-5-flash): names its maker in 100% of replies (2026-10-05) - [Kimi K2.7 Code](https://benchgecko.ai/model/kimi-k2-7-code): names its maker in 100% of replies (2026-10-05) - [MiniMax M3](https://benchgecko.ai/model/minimax-m3): names its maker in 100% of replies (2026-10-05) - [GPT-6 Luna Pro](https://benchgecko.ai/model/gpt-6-luna-pro): names its maker in 100% of replies (2026-10-04) - [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5): names its maker in 100% of replies (2026-10-04) - [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5): names its maker in 100% of replies (2026-10-04) - [Claude Fable 5.1](https://benchgecko.ai/model/claude-fable-5-1): names its maker in 100% of replies (2026-10-04) - [Claude Opus 5](https://benchgecko.ai/model/claude-opus-5): names its maker in 100% of replies (2026-10-04) - [Claude Opus 5 (Fast)](https://benchgecko.ai/model/claude-opus-5-fast): names its maker in 100% of replies (2026-10-04) - [Qwen3.8 Max Prime](https://benchgecko.ai/model/qwen3-8-max-prime): names its maker in 100% of replies (2026-10-04) - [Qwen3.8 Max (0902)](https://benchgecko.ai/model/qwen3-8-max-0902): names its maker in 100% of replies (2026-10-04) - [Qwen3.8 27B](https://benchgecko.ai/model/qwen3-8-27b): names its maker in 100% of replies (2026-10-04) - [Qwen3.8 2.4T A95B](https://benchgecko.ai/model/qwen3-8-2-4t-a95b): names its maker in 100% of replies (2026-10-04) - [GLM 5.3 Prime](https://benchgecko.ai/model/glm-5-3-prime): names its maker in 100% of replies (2026-10-04) - [GLM 5.3 FlashX](https://benchgecko.ai/model/glm-5-3-flashx): names its maker in 100% of replies (2026-10-04) - [GLM 5.3](https://benchgecko.ai/model/glm-5-3): names its maker in 100% of replies (2026-10-04) - [GLM 5.1](https://benchgecko.ai/model/glm-5-1): names its maker in 100% of replies (2026-10-04) - [Grok 4.7](https://benchgecko.ai/model/grok-4-7): names its maker in 100% of replies (2026-10-04) - [Grok 4.6](https://benchgecko.ai/model/grok-4-6): names its maker in 100% of replies (2026-10-04) - [Grok 4.5](https://benchgecko.ai/model/grok-4-5): names its maker in 100% of replies (2026-10-04) - [Grok 4.3](https://benchgecko.ai/model/grok-4-3): names its maker in 100% of replies (2026-10-04) - [Grok 4.20](https://benchgecko.ai/model/grok-4-20): names its maker in 100% of replies (2026-10-04) - [DeepSeek V4 Flash](https://benchgecko.ai/model/deepseek-v4-flash): names its maker in 100% of replies (2026-10-04) - [DeepSeek V4 Pro](https://benchgecko.ai/model/deepseek-v4-pro): names its maker in 100% of replies (2026-10-04) - [Gemini 3.8 Flash](https://benchgecko.ai/model/gemini-3-8-flash): names its maker in 100% of replies (2026-10-04) - [Gemini 3.7 Flash](https://benchgecko.ai/model/gemini-3-7-flash): names its maker in 100% of replies (2026-10-04) - [Gemini 3.6 Flash](https://benchgecko.ai/model/gemini-3-6-flash): names its maker in 100% of replies (2026-10-04) - [Gemini 3.5 Flash Lite](https://benchgecko.ai/model/gemini-3-5-flash-lite): names its maker in 100% of replies (2026-10-04) - [Gemma 4 26B A4B ](https://benchgecko.ai/model/gemma-4-26b-a4b-it): names its maker in 100% of replies (2026-10-04) - [Kimi K2.6](https://benchgecko.ai/model/kimi-k2-6): names its maker in 100% of replies (2026-10-04) - [Kimi K2.5](https://benchgecko.ai/model/kimi-k2-5): names its maker in 100% of replies (2026-10-04) - [Kimi K2 Thinking](https://benchgecko.ai/model/kimi-k2-thinking): names its maker in 100% of replies (2026-10-04) - [Kimi K2 0905](https://benchgecko.ai/model/kimi-k2-0905): names its maker in 100% of replies (2026-10-04) - [MiniMax M2.7](https://benchgecko.ai/model/minimax-m2-7): names its maker in 100% of replies (2026-10-04) - [GPT-6 Luna](https://benchgecko.ai/model/gpt-6-luna): names its maker in 100% of replies (2026-10-04) - [GLM 5.3 Flash](https://benchgecko.ai/model/glm-5-3-flash): names its maker in 100% of replies (2026-10-04) - [Qwen3.8 Flash](https://benchgecko.ai/model/qwen3-8-flash): names its maker in 100% of replies (2026-10-04) - [Kimi K3](https://benchgecko.ai/model/kimi-k3): names its maker in 100% of replies (2026-10-04) - [GPT-6.1 Sol Pro](https://benchgecko.ai/model/gpt-6-1-sol-pro): names its maker in 100% of replies (2026-10-04) - [GPT-6.1 Sol](https://benchgecko.ai/model/gpt-6-1-sol): names its maker in 100% of replies (2026-10-04) ### World Map grid 6° How well does the model draw the world map from memory? Metric: Accuracy of land or water answers on a 6 degree grid against Natural Earth 1:110m. As of 2026-10-03. Page: https://benchgecko.ai/gecko-tests/world-map. Data: https://benchgecko.ai/api/v1/export/world-map - [Gemini 3.8 Flash](https://benchgecko.ai/model/gemini-3-8-flash): 97.4% accuracy, land recall 97.2%, water recall 97.4% (2026-10-03) - [GPT-6 Luna](https://benchgecko.ai/model/gpt-6-luna): 96.2% accuracy, land recall 94.6%, water recall 95.7% (2026-10-03) - [MiniMax M3](https://benchgecko.ai/model/minimax-m3): 92.8% accuracy, land recall 89.9%, water recall 94% (2026-10-03) - [GLM 5.3 Flash](https://benchgecko.ai/model/glm-5-3-flash): 89.7% accuracy, land recall 91.4%, water recall 88.7% (2026-10-03) - [Llama 4 Maverick](https://benchgecko.ai/model/llama-4-maverick): 79.2% accuracy, land recall 60.2%, water recall 89% (2026-10-03) ### Censorship Index v2 How often does the model refuse legitimate questions? Metric: Refusal rate: refused plus redirected replies over valid replies, 40 questions in 8 categories. As of 2026-10-05. Page: https://benchgecko.ai/gecko-tests/censorship-index. Data: https://benchgecko.ai/api/lab/censorship-index - [DeepSeek V4.1 Flash](https://benchgecko.ai/model/deepseek-v4-1-flash): refusal rate 5.6%, answered 32/36 (2026-10-04) - [Claude Opus 5](https://benchgecko.ai/model/claude-opus-5): refusal rate 2.6%, answered 33/39 (2026-10-05) - [GPT-6.1 Sol Pro](https://benchgecko.ai/model/gpt-6-1-sol-pro): refusal rate 2.5%, answered 36/40 (2026-10-05) - [GPT-6.1 Sol](https://benchgecko.ai/model/gpt-6-1-sol): refusal rate 2.5%, answered 36/40 (2026-10-05) - [GPT-6 Luna Pro](https://benchgecko.ai/model/gpt-6-luna-pro): refusal rate 2.5%, answered 37/40 (2026-10-05) - [GPT-6 Luna](https://benchgecko.ai/model/gpt-6-luna): refusal rate 2.5%, answered 37/40 (2026-10-05) - [GPT-6 Sol](https://benchgecko.ai/model/gpt-6-sol): refusal rate 0%, answered 38/40 (2026-10-05) - [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5): refusal rate 0%, answered 38/40 (2026-10-05) - [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5): refusal rate 0%, answered 35/39 (2026-10-05) ### Knowledge Horizon v1 Where does the model's knowledge of world events actually stop? Metric: Last month whose three month rolling accuracy is at least half of the model's 2024 accuracy. As of 2026-10-04. Page: https://benchgecko.ai/gecko-tests/knowledge-horizon. Data: https://benchgecko.ai/api/lab/knowledge-horizon - [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5): measured horizon 2026-05, claimed cutoff not published, released 2026-09-28 (2026-10-04) - [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5): measured horizon 2026-05, claimed cutoff not published, released 2026-09-22 (2026-10-04) - [GPT-6 Luna Pro](https://benchgecko.ai/model/gpt-6-luna-pro): measured horizon 2026-04, claimed cutoff not published, released 2026-09-22 (2026-10-04) - [GPT-6 Sol](https://benchgecko.ai/model/gpt-6-sol): measured horizon 2026-04, claimed cutoff not published, released 2026-09-22 (2026-10-04) - [GPT-6 Luna](https://benchgecko.ai/model/gpt-6-luna): measured horizon 2026-04, claimed cutoff not published, released 2026-09-22 (2026-10-04) - [GPT-6.1 Sol Pro](https://benchgecko.ai/model/gpt-6-1-sol-pro): measured horizon 2026-03, claimed cutoff not published, released 2026-09-29 (2026-10-04) - [GPT-6.1 Sol](https://benchgecko.ai/model/gpt-6-1-sol): measured horizon 2026-03, claimed cutoff not published, released 2026-09-29 (2026-10-04) ### Tokenizer Tax v1 How many more tokens does the same text cost outside English? Metric: Average extra tokens across eight languages versus English, from the model's own token counts. As of 2026-10-05. Page: https://benchgecko.ai/gecko-tests/tokenizer-tax. Data: https://benchgecko.ai/api/lab/tokenizer-tax - [DeepSeek V4 Flash](https://benchgecko.ai/model/deepseek-v4-flash): +7.4% tokens outside English; ja 1.09x, ko 1.27x (2026-10-04) - [DeepSeek V4 Pro 0813](https://benchgecko.ai/model/deepseek-v4-pro-0813): +17.1% tokens outside English; ja 1.09x, ko 1.27x (2026-10-04) - [MiniMax M2.7](https://benchgecko.ai/model/minimax-m2-7): +34.5% tokens outside English; ja 1.2x, ko 1.49x (2026-10-04) - [Qwen3.8 Max Prime](https://benchgecko.ai/model/qwen3-8-max-prime): +35.4% tokens outside English; ja 1.44x, ko 1.41x (2026-10-04) - [Qwen3.8 Max (0902)](https://benchgecko.ai/model/qwen3-8-max-0902): +35.4% tokens outside English; ja 1.44x, ko 1.41x (2026-10-04) - [Qwen3.8 Flash](https://benchgecko.ai/model/qwen3-8-flash): +35.4% tokens outside English; ja 1.44x, ko 1.41x (2026-10-04) - [MiniMax M3](https://benchgecko.ai/model/minimax-m3): +36.4% tokens outside English; ja 1.2x, ko 1.52x (2026-10-05) - [MiniMax M2-her](https://benchgecko.ai/model/minimax-m2-her): +36.4% tokens outside English; ja 1.2x, ko 1.52x (2026-10-04) - [MiniMax M2.1](https://benchgecko.ai/model/minimax-m2-1): +36.4% tokens outside English; ja 1.2x, ko 1.52x (2026-10-04) - [MiniMax M2](https://benchgecko.ai/model/minimax-m2): +36.4% tokens outside English; ja 1.2x, ko 1.52x (2026-10-04) - [Qwen3.8 2.4T A95B](https://benchgecko.ai/model/qwen3-8-2-4t-a95b): +38.9% tokens outside English; ja 1.44x, ko 1.41x (2026-10-04) - [Qwen3.8 27B](https://benchgecko.ai/model/qwen3-8-27b): +40.5% tokens outside English; ja 1.85x, ko 1.41x (2026-10-04) - [Llama 4 Scout](https://benchgecko.ai/model/llama-4-scout): +42.8% tokens outside English; ja 1.63x, ko 1.51x (2026-10-04) - [Grok Build 0.1](https://benchgecko.ai/model/grok-build-0-1): +43.5% tokens outside English; ja 1.56x, ko 1.77x (2026-10-05) - [Grok 4.3](https://benchgecko.ai/model/grok-4-3): +43.5% tokens outside English; ja 1.56x, ko 1.77x (2026-10-04) - [Grok 4.20](https://benchgecko.ai/model/grok-4-20): +43.5% tokens outside English; ja 1.56x, ko 1.77x (2026-10-04) - [Gemini 3.8 Flash](https://benchgecko.ai/model/gemini-3-8-flash): +43.6% tokens outside English; ja 1.48x, ko 1.64x (2026-10-04) - [Gemini 3.7 Flash](https://benchgecko.ai/model/gemini-3-7-flash): +43.6% tokens outside English; ja 1.48x, ko 1.64x (2026-10-04) - [Gemini 3.6 Flash](https://benchgecko.ai/model/gemini-3-6-flash): +43.6% tokens outside English; ja 1.48x, ko 1.64x (2026-10-04) - [Gemini 3.5 Flash Lite](https://benchgecko.ai/model/gemini-3-5-flash-lite): +43.6% tokens outside English; ja 1.48x, ko 1.64x (2026-10-04) - [Gemma 4 26B A4B ](https://benchgecko.ai/model/gemma-4-26b-a4b-it): +43.6% tokens outside English; ja 1.48x, ko 1.64x (2026-10-04) - [Llama 4 Maverick](https://benchgecko.ai/model/llama-4-maverick): +44.1% tokens outside English; ja 1.62x, ko 1.64x (2026-10-04) - [Grok 4.7](https://benchgecko.ai/model/grok-4-7): +50.4% tokens outside English; ja 1.43x, ko 2x (2026-10-04) - [Grok 4.6](https://benchgecko.ai/model/grok-4-6): +50.4% tokens outside English; ja 1.43x, ko 2x (2026-10-04) - [Grok 4.5](https://benchgecko.ai/model/grok-4-5): +50.4% tokens outside English; ja 1.43x, ko 2x (2026-10-04) - [Mistral Medium 3.5](https://benchgecko.ai/model/mistral-medium-3-5): +53.3% tokens outside English; ja 2.05x, ko 1.47x (2026-10-05) - [GPT-6 Sol](https://benchgecko.ai/model/gpt-6-sol): +53.3% tokens outside English; ja 2.08x, ko 1.74x (2026-10-04) - [Mistral Small 4](https://benchgecko.ai/model/mistral-small-2603): +53.3% tokens outside English; ja 2.05x, ko 1.47x (2026-10-04) - [Devstral 2 2512](https://benchgecko.ai/model/devstral-2512): +53.3% tokens outside English; ja 2.05x, ko 1.47x (2026-10-04) - [Ministral 3 14B 2512](https://benchgecko.ai/model/ministral-14b-2512): +53.3% tokens outside English; ja 2.05x, ko 1.47x (2026-10-04) - [Ministral 3 3B 2512](https://benchgecko.ai/model/ministral-3b-2512): +53.3% tokens outside English; ja 2.05x, ko 1.47x (2026-10-04) - [GPT-6 Luna](https://benchgecko.ai/model/gpt-6-luna): +56.1% tokens outside English; ja 2.14x, ko 1.78x (2026-10-04) - [GPT-6 Luna Pro](https://benchgecko.ai/model/gpt-6-luna-pro): +56.5% tokens outside English; ja 2.14x, ko 1.79x (2026-10-04) - [GPT-6.1 Sol Pro](https://benchgecko.ai/model/gpt-6-1-sol-pro): +63.8% tokens outside English; ja 2.11x, ko 1.76x (2026-10-05) - [MiniMax M2.5](https://benchgecko.ai/model/minimax-m2-5): +63.9% tokens outside English; ja 2.29x, ko 1.54x (2026-10-04) - [DeepSeek V4.1 Flash](https://benchgecko.ai/model/deepseek-v4-1-flash): +68% tokens outside English; ja 1.79x, ko 2.07x (2026-10-04) - [DeepSeek V4 Flash 0731](https://benchgecko.ai/model/deepseek-v4-flash-0731): +68% tokens outside English; ja 1.79x, ko 2.07x (2026-10-04) - [DeepSeek V4 Pro](https://benchgecko.ai/model/deepseek-v4-pro): +68% tokens outside English; ja 1.79x, ko 2.07x (2026-10-04) - [GLM 5.3 Prime](https://benchgecko.ai/model/glm-5-3-prime): +69.6% tokens outside English; ja 1.89x, ko 2.4x (2026-10-04) - [GLM 5.3 FlashX](https://benchgecko.ai/model/glm-5-3-flashx): +69.6% tokens outside English; ja 1.89x, ko 2.4x (2026-10-04) - [GLM 5.3 Flash](https://benchgecko.ai/model/glm-5-3-flash): +69.6% tokens outside English; ja 1.89x, ko 2.4x (2026-10-04) - [GLM 5.1](https://benchgecko.ai/model/glm-5-1): +70.9% tokens outside English; ja 1.89x, ko 2.4x (2026-10-04) - [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5): +72.6% tokens outside English; ja 1.58x, ko 2.11x (2026-10-04) - [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5): +72.6% tokens outside English; ja 1.58x, ko 2.11x (2026-10-04) - [Claude Fable 5.1](https://benchgecko.ai/model/claude-fable-5-1): +72.6% tokens outside English; ja 1.58x, ko 2.11x (2026-10-04) - [Claude Opus 5](https://benchgecko.ai/model/claude-opus-5): +72.6% tokens outside English; ja 1.58x, ko 2.11x (2026-10-04) - [Claude Opus 5 (Fast)](https://benchgecko.ai/model/claude-opus-5-fast): +72.6% tokens outside English; ja 1.58x, ko 2.11x (2026-10-04) - [GLM 5.3](https://benchgecko.ai/model/glm-5-3): +90.1% tokens outside English; ja 2.11x, ko 2.75x (2026-10-04) - [Kimi K2 0905](https://benchgecko.ai/model/kimi-k2-0905): +91.9% tokens outside English; ja 2.13x, ko 2.55x (2026-10-04) - [Kimi K2.7 Code](https://benchgecko.ai/model/kimi-k2-7-code): +92% tokens outside English; ja 2.13x, ko 2.55x (2026-10-05) - [Kimi K2.6](https://benchgecko.ai/model/kimi-k2-6): +92% tokens outside English; ja 2.13x, ko 2.55x (2026-10-04) - [Kimi K2 Thinking](https://benchgecko.ai/model/kimi-k2-thinking): +92.3% tokens outside English; ja 2.13x, ko 2.55x (2026-10-04) - [Kimi K3](https://benchgecko.ai/model/kimi-k3): +93% tokens outside English; ja 2.15x, ko 2.55x (2026-10-04) - [Kimi K2.5](https://benchgecko.ai/model/kimi-k2-5): +99.5% tokens outside English; ja 2.28x, ko 2.55x (2026-10-04) ### Same Model, Different Host Do providers serving the same open model give the same quality? Metric: Share of 30 exact-answer tasks right per provider; flagged when more than 15 points below the best host. As of n/a. Page: https://benchgecko.ai/gecko-tests/same-model-different-host. Data: https://benchgecko.ai/api/lab/same-model-different-host - First results coming. ### Model Drift Index v1 Do models quietly change behind the same name? Metric: Weekly run of 22 fixed probes compared with the previous run (accuracy, refusals, self identification). As of 2026-10-05. Page: https://benchgecko.ai/gecko-tests/model-drift-index. Data: https://benchgecko.ai/api/lab/model-drift - [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5): baseline, exact answers 100%, refusals 0/8 (2026-10-05) - [DeepSeek V4 Pro 0813](https://benchgecko.ai/model/deepseek-v4-pro-0813): baseline, exact answers 100%, refusals 0/7 (2026-10-05) - [GLM 5.3 Prime](https://benchgecko.ai/model/glm-5-3-prime): baseline, exact answers 80%, refusals 0/8 (2026-10-05) - [GPT-6.1 Sol Pro](https://benchgecko.ai/model/gpt-6-1-sol-pro): baseline, exact answers 100%, refusals 0/8 (2026-10-05) - [Grok 4.7](https://benchgecko.ai/model/grok-4-7): baseline, exact answers 100%, refusals 0/8 (2026-10-05) - [Kimi K3](https://benchgecko.ai/model/kimi-k3): baseline, exact answers 100%, refusals 0/8 (2026-10-05) - [Llama 4 Scout](https://benchgecko.ai/model/llama-4-scout): baseline, exact answers 20%, refusals 0/8 (2026-10-05) - [Mistral Medium 3.5](https://benchgecko.ai/model/mistral-medium-3-5): baseline, exact answers 20%, refusals 0/8 (2026-10-05) - [Qwen3.8 Max Prime](https://benchgecko.ai/model/qwen3-8-max-prime): baseline, exact answers 100%, refusals 0/8 (2026-10-05) ### Gecko Scorecard Each test gives a 0 to 100 value per model; grades are relative (A = top 20% of models that took the test, then B, C, D, E). Gecko Score = average percentile rank across the model's tests, shown once a model has at least 3 tests. - 1. Gemini 3.8 Flash: Gecko Score 77 · who-are-you B (Knows who made it) · world-map A (97.4% of the map right) · tokenizer-tax B (44% more tokens outside English) - 2. MiniMax M3: Gecko Score 67 · who-are-you B (Knows who made it) · world-map C (92.8% of the map right) · tokenizer-tax A (36% more tokens outside English) - 3. Claude Opus 5.5: Gecko Score 65 · who-are-you B (Knows who made it) · censorship-index A (Answers almost everything) · knowledge-horizon A (Knows news up to May 2026) · tokenizer-tax E (73% more tokens outside English) - 4. Claude Sonnet 5.5: Gecko Score 65 · who-are-you B (Knows who made it) · censorship-index A (Answers almost everything) · knowledge-horizon A (Knows news up to May 2026) · tokenizer-tax E (73% more tokens outside English) - 5. GPT-6 Luna: Gecko Score 55 · who-are-you B (Knows who made it) · world-map B (96.2% of the map right) · censorship-index C (Answers almost everything) · knowledge-horizon C (Knows news up to Apr 2026) · tokenizer-tax C (56% more tokens outside English) - 6. GPT-6 Sol: Gecko Score 52 · who-are-you D (Knows who made it) · censorship-index A (Answers almost everything) · knowledge-horizon C (Knows news up to Apr 2026) · tokenizer-tax C (53% more tokens outside English) - 7. GPT-6 Luna Pro: Gecko Score 49 · who-are-you B (Knows who made it) · censorship-index C (Answers almost everything) · knowledge-horizon C (Knows news up to Apr 2026) · tokenizer-tax D (57% more tokens outside English) - 8. GPT-6.1 Sol Pro: Gecko Score 39 · who-are-you B (Knows who made it) · censorship-index C (Answers almost everything) · knowledge-horizon E (Knows news up to Mar 2026) · tokenizer-tax D (64% more tokens outside English) - 9. GLM 5.3 Flash: Gecko Score 39 · who-are-you B (Knows who made it) · world-map D (89.7% of the map right) · tokenizer-tax D (70% more tokens outside English) - 10. GPT-6.1 Sol: Gecko Score 39 · who-are-you B (Knows who made it) · censorship-index C (Answers almost everything) · knowledge-horizon E (Knows news up to Mar 2026) - 11. Claude Opus 5: Gecko Score 31 · who-are-you B (Knows who made it) · censorship-index E (Answers almost everything) · tokenizer-tax E (73% more tokens outside English) - 12. Llama 4 Maverick: Gecko Score 20 · who-are-you E (Sometimes says it is Google or OpenAI) · world-map E (79.2% of the map right) · tokenizer-tax B (44% more tokens outside English) - 13. DeepSeek V4.1 Flash: Gecko Score 12 · who-are-you E (Sometimes says it is OpenAI) · censorship-index E (Refuses or dodges 6%) · tokenizer-tax D (68% more tokens outside English) ## Models and benchmarks - [All models](https://benchgecko.ai/models): leaderboard by BenchGecko score (normalized average of public benchmark scores), with price, context and license. - [Model pages](https://benchgecko.ai/model/claude-opus-5-5): /model/ · score, rank, list price, providers, benchmarks with sources, Gecko Tests grades, price history. - [Benchmarks](https://benchgecko.ai/benchmarks): every tracked benchmark and its ranked models. - [Compare](https://benchgecko.ai/compare): /compare/-vs- side by side. - [Methodology](https://benchgecko.ai/methodology): sources, normalization and refresh schedule. Top 10 by BenchGecko score (as of 2026-10-05): - 1. [GPT-5.5 Pro](https://benchgecko.ai/model/gpt-5-5-pro) (OpenAI): score 99.9, list price $30.00 input / $180.00 output per 1M tokens - 2. [Claude Mythos Preview](https://benchgecko.ai/model/claude-mythos-preview) (Anthropic): score 99.8, list price n/a input / n/a output per 1M tokens - 3. [Claude Opus 5.5](https://benchgecko.ai/model/claude-opus-5-5) (Anthropic): score 96.8, list price $4.00 input / $20.00 output per 1M tokens - 4. [GPT-6 Astra](https://benchgecko.ai/model/gpt-6-astra) (OpenAI): score 96.3, list price $10.00 input / $50.00 output per 1M tokens - 5. [DeepSeek V3.2 Speciale](https://benchgecko.ai/model/deepseek-v3-2-speciale) (DeepSeek): score 95.2, list price $0.40 input / $1.20 output per 1M tokens - 6. [Step 3.5 Flash](https://benchgecko.ai/model/step-3-5-flash) (stepfun): score 89.5, list price $0.10 input / $0.30 output per 1M tokens - 7. [Claude Sonnet 5.5](https://benchgecko.ai/model/claude-sonnet-5-5) (Anthropic): score 89.2, list price $2.00 input / $10.00 output per 1M tokens - 8. [GPT-5 Chat](https://benchgecko.ai/model/gpt-5-chat) (OpenAI): score 89, list price $1.25 input / $10.00 output per 1M tokens - 9. [Claude Fable 5.1](https://benchgecko.ai/model/claude-fable-5-1) (Anthropic): score 88.2, list price $10.00 input / $50.00 output per 1M tokens - 10. [Qwen2.5 72B Instruct Abliterated](https://benchgecko.ai/model/huihui-ai-qwen25-72b-instruct-abliterated) (HuiHui AI): score 87.2, list price n/a input / n/a output per 1M tokens ## Pricing - [AI pricing](https://benchgecko.ai/pricing): every model and every provider, filtered by use case and budget. - [Price per provider](https://benchgecko.ai/pricing/arbitrage): the same model at every provider, cheapest first; /pricing/arbitrage/ per model. - [Price drops](https://benchgecko.ai/pricing/drops): recorded price changes. - [Free tiers](https://benchgecko.ai/pricing/free): free and free-quota APIs. ## Companies, mindshare and infrastructure - [AI economy](https://benchgecko.ai/economy): valuations, funding, revenue and the AI Bubble Index; each figure has an as-of date and a source. - [Company pages](https://benchgecko.ai/economy/companies): /economy/company/. - [Mindshare](https://benchgecko.ai/mindshare): daily attention share from Hacker News, Wikipedia pageviews and GitHub stars. Top model mindshare (as of 2026-10-04): GPT-6 24.1% · Qwen 22.7% · GLM 14.2% · GPT-5 7.0% · Gemini 3 5.9% - [Hardware](https://benchgecko.ai/hardware): AI chips with specs and daily GPU rental prices (Vast.ai marketplace median per GPU hour). - [Providers](https://benchgecko.ai/providers): every API provider and its models. - [Status](https://benchgecko.ai/status): provider uptime pings. ## Top 100 models by BenchGecko score - 1. GPT-5.5 Pro (OpenAI): score 99.9 · $30.00 in / $180.00 out per 1M · https://benchgecko.ai/model/gpt-5-5-pro - 2. Claude Mythos Preview (Anthropic): score 99.8 · n/a in / n/a out per 1M · https://benchgecko.ai/model/claude-mythos-preview - 3. Claude Opus 5.5 (Anthropic): score 96.8 · $4.00 in / $20.00 out per 1M · https://benchgecko.ai/model/claude-opus-5-5 - 4. GPT-6 Astra (OpenAI): score 96.3 · $10.00 in / $50.00 out per 1M · https://benchgecko.ai/model/gpt-6-astra - 5. DeepSeek V3.2 Speciale (DeepSeek): score 95.2 · $0.40 in / $1.20 out per 1M · https://benchgecko.ai/model/deepseek-v3-2-speciale - 6. Step 3.5 Flash (stepfun): score 89.5 · $0.10 in / $0.30 out per 1M · https://benchgecko.ai/model/step-3-5-flash - 7. Claude Sonnet 5.5 (Anthropic): score 89.2 · $2.00 in / $10.00 out per 1M · https://benchgecko.ai/model/claude-sonnet-5-5 - 8. GPT-5 Chat (OpenAI): score 89 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gpt-5-chat - 9. Claude Fable 5.1 (Anthropic): score 88.2 · $10.00 in / $50.00 out per 1M · https://benchgecko.ai/model/claude-fable-5-1 - 10. Qwen2.5 72B Instruct Abliterated (HuiHui AI): score 87.2 · n/a in / n/a out per 1M · https://benchgecko.ai/model/huihui-ai-qwen25-72b-instruct-abliterated - 11. Claude Opus 5 (Anthropic): score 86.5 · $5.00 in / $25.00 out per 1M · https://benchgecko.ai/model/claude-opus-5 - 12. DeepSeek-V2 (MoE-236B, May 2024) (DeepSeek): score 85.9 · n/a in / n/a out per 1M · https://benchgecko.ai/model/deepseek-v2-moe-236b-may-2024 - 13. GPT-5.1-Codex-Max (OpenAI): score 85.6 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gpt-5-1-codex-max - 14. Claude Fable 5 (Anthropic): score 83.6 · $10.00 in / $50.00 out per 1M · https://benchgecko.ai/model/claude-fable-5 - 15. Claude Opus 4.6 (Fast) (Anthropic): score 83.3 · $30.00 in / $150.00 out per 1M · https://benchgecko.ai/model/claude-opus-4-6-fast - 16. PaLM 2-L (Unknown): score 81.8 · n/a in / n/a out per 1M · https://benchgecko.ai/model/palm-2-l - 17. MiMo-V2-Flash (xiaomi): score 81.7 · $0.0900 in / $0.29 out per 1M · https://benchgecko.ai/model/mimo-v2-flash - 18. GPT-5.6 Sol (OpenAI): score 80.9 · $2.00 in / $10.00 out per 1M · https://benchgecko.ai/model/gpt-5-6-sol - 19. GPT-5.4 Pro (OpenAI): score 80.7 · $30.00 in / $180.00 out per 1M · https://benchgecko.ai/model/gpt-5-4-pro - 20. GPT-5.2-Codex (OpenAI): score 80.7 · $1.75 in / $14.00 out per 1M · https://benchgecko.ai/model/gpt-5-2-codex - 21. phi-3-small 7.4B (Microsoft): score 79.2 · n/a in / n/a out per 1M · https://benchgecko.ai/model/phi-3-small-7-4b - 22. GPT-5.1-Codex (OpenAI): score 77.7 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gpt-5-1-codex - 23. Grok Build 0.1 (xAI): score 77.1 · $1.00 in / $2.00 out per 1M · https://benchgecko.ai/model/grok-build-0-1 - 24. Qwen2 7B Instruct (Alibaba): score 76.5 · n/a in / n/a out per 1M · https://benchgecko.ai/model/qwen-qwen2-7b-instruct - 25. GPT-5.6 Terra (OpenAI): score 76.5 · $2.00 in / $12.00 out per 1M · https://benchgecko.ai/model/gpt-5-6-terra - 26. gpt-oss-120b (free) (OpenAI): score 74.2 · free in / free out per 1M · https://benchgecko.ai/model/gpt-oss-120b-free - 27. DeepSeek R1 Distill Qwen 14B (DeepSeek): score 73.7 · n/a in / n/a out per 1M · https://benchgecko.ai/model/deepseek-ai-deepseek-r1-distill-qwen-14b - 28. Gemini 2.5 Pro Preview 06-05 (Google DeepMind): score 73.5 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gemini-2-5-pro-preview - 29. Muse Spark (Unknown): score 73.3 · n/a in / n/a out per 1M · https://benchgecko.ai/model/muse-spark - 30. Hermes 3 70B Instruct (nousresearch): score 73.2 · $0.70 in / $0.70 out per 1M · https://benchgecko.ai/model/hermes-3-llama-3-1-70b - 31. MiniMax M2 (minimax): score 72.4 · $0.30 in / $1.20 out per 1M · https://benchgecko.ai/model/minimax-m2 - 32. GLM 4.5 (z-ai): score 72.3 · $0.60 in / $2.20 out per 1M · https://benchgecko.ai/model/glm-4-5 - 33. Claude Opus 4.8 (Anthropic): score 71.8 · $5.00 in / $25.00 out per 1M · https://benchgecko.ai/model/claude-opus-4-8 - 34. Qwen2.5 Coder 32B Instruct (Alibaba Qwen): score 71.4 · $0.66 in / $1.00 out per 1M · https://benchgecko.ai/model/qwen-2-5-coder-32b-instruct - 35. Qwen2.5 14B Instruct (Alibaba): score 71 · n/a in / n/a out per 1M · https://benchgecko.ai/model/qwen-qwen25-14b-instruct - 36. GPT-5.2 Pro (OpenAI): score 70.2 · $21.00 in / $168.00 out per 1M · https://benchgecko.ai/model/gpt-5-2-pro - 37. Qwen3.7 Max (Alibaba Qwen): score 69.8 · $1.25 in / $3.75 out per 1M · https://benchgecko.ai/model/qwen3-7-max - 38. Qwen3.5 397B A17B (Alibaba Qwen): score 69.7 · $0.55 in / $3.50 out per 1M · https://benchgecko.ai/model/qwen3-5-397b-a17b - 39. phi-3-medium 14B (Microsoft): score 69.2 · n/a in / n/a out per 1M · https://benchgecko.ai/model/phi-3-medium-14b - 40. Kimi K3 (moonshotai): score 68.9 · $0.67 in / $14.00 out per 1M · https://benchgecko.ai/model/kimi-k3 - 41. Gemini 3 Pro (Google DeepMind): score 68.7 · n/a in / n/a out per 1M · https://benchgecko.ai/model/gemini-3-pro - 42. Claude Instant (Anthropic): score 68.7 · n/a in / n/a out per 1M · https://benchgecko.ai/model/claude-instant - 43. DeepSeek V4 Flash (DeepSeek): score 68.6 · $0.0250 in / $1.28 out per 1M · https://benchgecko.ai/model/deepseek-v4-flash - 44. Mixtral 8x7B Instruct (Mistral AI): score 68.2 · $0.54 in / $0.54 out per 1M · https://benchgecko.ai/model/mixtral-8x7b-instruct - 45. GPT-5.3-Codex (OpenAI): score 68.1 · $1.75 in / $14.00 out per 1M · https://benchgecko.ai/model/gpt-5-3-codex - 46. PaLM 2-M (Unknown): score 68.1 · n/a in / n/a out per 1M · https://benchgecko.ai/model/palm-2-m - 47. Grok 3 Beta (xAI): score 67.9 · $3.00 in / $15.00 out per 1M · https://benchgecko.ai/model/grok-3-beta - 48. Gemini 3.8 Flash (Google DeepMind): score 67.7 · $0.75 in / $3.75 out per 1M · https://benchgecko.ai/model/gemini-3-8-flash - 49. DeepSeek V4 Pro 0813 (DeepSeek): score 66.8 · $0.55 in / $5.00 out per 1M · https://benchgecko.ai/model/deepseek-v4-pro-0813 - 50. Gemini 3.7 Flash (Google DeepMind): score 66.4 · $0.75 in / $3.75 out per 1M · https://benchgecko.ai/model/gemini-3-7-flash - 51. Claude Opus 4.6 (Anthropic): score 66.4 · $5.00 in / $25.00 out per 1M · https://benchgecko.ai/model/claude-opus-4-6 - 52. Claude Sonnet 5 (Anthropic): score 66.3 · $2.00 in / $10.00 out per 1M · https://benchgecko.ai/model/claude-sonnet-5 - 53. Grok 4.6 (xAI): score 66.2 · $2.00 in / $6.00 out per 1M · https://benchgecko.ai/model/grok-4-6 - 54. GPT-5.4 (OpenAI): score 66 · $2.50 in / $15.00 out per 1M · https://benchgecko.ai/model/gpt-5-4 - 55. GPT-5.5 (OpenAI): score 65.7 · $5.00 in / $30.00 out per 1M · https://benchgecko.ai/model/gpt-5-5 - 56. Claude Opus 4.7 (Anthropic): score 65.6 · $5.00 in / $25.00 out per 1M · https://benchgecko.ai/model/claude-opus-4-7 - 57. Gemini 3.1 Pro Preview (Google DeepMind): score 65 · $2.00 in / $12.00 out per 1M · https://benchgecko.ai/model/gemini-3-1-pro-preview - 58. Qwen2 VL 7B Instruct (Alibaba): score 64.9 · n/a in / n/a out per 1M · https://benchgecko.ai/model/qwen-qwen2-vl-7b-instruct - 59. GPT-5.6 Luna (OpenAI): score 64.8 · $0.20 in / $1.20 out per 1M · https://benchgecko.ai/model/gpt-5-6-luna - 60. GLM 5.2 (z-ai): score 64.7 · $1.00 in / $4.00 out per 1M · https://benchgecko.ai/model/glm-5-2 - 61. Falcon 2 11B (TII): score 64.3 · n/a in / n/a out per 1M · https://benchgecko.ai/model/falcon-2-11b - 62. Qwen2.5 Coder 14B Instruct (Alibaba): score 64.1 · n/a in / n/a out per 1M · https://benchgecko.ai/model/qwen-qwen25-coder-14b-instruct - 63. GLM 5 (z-ai): score 64.1 · $0.60 in / $1.92 out per 1M · https://benchgecko.ai/model/glm-5 - 64. Meta Llama 3 8B (Meta): score 64 · n/a in / n/a out per 1M · https://benchgecko.ai/model/meta-llama-meta-llama-3-8b - 65. GPT-4 (older v0314) (OpenAI): score 63.8 · $30.00 in / $60.00 out per 1M · https://benchgecko.ai/model/gpt-4-0314 - 66. GLM 5.1 (z-ai): score 63.8 · $0.97 in / $3.04 out per 1M · https://benchgecko.ai/model/glm-5-1 - 67. Kimi K2.7 Code (moonshotai): score 63.6 · $0.61 in / $3.07 out per 1M · https://benchgecko.ai/model/kimi-k2-7-code - 68. Mixtral 8x7B (Mistral AI): score 63.6 · n/a in / n/a out per 1M · https://benchgecko.ai/model/mixtral-8x7b - 69. Qwen3 30B A3B Thinking 2507 (Alibaba Qwen): score 63.5 · $0.20 in / $2.40 out per 1M · https://benchgecko.ai/model/qwen3-30b-a3b-thinking-2507 - 70. GPT-5.1-Codex-Mini (OpenAI): score 63.3 · $0.25 in / $2.00 out per 1M · https://benchgecko.ai/model/gpt-5-1-codex-mini - 71. GLM 5.3 (z-ai): score 63.1 · $1.40 in / $4.40 out per 1M · https://benchgecko.ai/model/glm-5-3 - 72. Qwen3.6 Plus (Alibaba Qwen): score 62.8 · $0.33 in / $1.95 out per 1M · https://benchgecko.ai/model/qwen3-6-plus - 73. Grok 4.5 (xAI): score 62.4 · $2.00 in / $6.00 out per 1M · https://benchgecko.ai/model/grok-4-5 - 74. Qwen3.6 27B (Alibaba Qwen): score 62.2 · $0.32 in / $2.70 out per 1M · https://benchgecko.ai/model/qwen3-6-27b - 75. PaLM 2-S (Unknown): score 61.7 · n/a in / n/a out per 1M · https://benchgecko.ai/model/palm-2-s - 76. Grok 4.3 (xAI): score 61.7 · $1.25 in / $2.50 out per 1M · https://benchgecko.ai/model/grok-4-3 - 77. phi-3-mini 3.8B (Microsoft): score 61.5 · n/a in / n/a out per 1M · https://benchgecko.ai/model/phi-3-mini-3-8b - 78. WizardLM-2 8x22B (Microsoft): score 61.4 · $0.62 in / $0.62 out per 1M · https://benchgecko.ai/model/wizardlm-2-8x22b - 79. Gemini 3.5 Flash (Google DeepMind): score 61.3 · $1.50 in / $9.00 out per 1M · https://benchgecko.ai/model/gemini-3-5-flash - 80. gpt-oss-20b (free) (OpenAI): score 61 · free in / free out per 1M · https://benchgecko.ai/model/gpt-oss-20b-free - 81. DeepSeek V4 Pro (DeepSeek): score 60.8 · $0.21 in / $0.42 out per 1M · https://benchgecko.ai/model/deepseek-v4-pro - 82. o3 Pro (OpenAI): score 60.8 · $20.00 in / $80.00 out per 1M · https://benchgecko.ai/model/o3-pro - 83. Gemini 3.6 Flash (Google DeepMind): score 60.6 · $0.75 in / $3.75 out per 1M · https://benchgecko.ai/model/gemini-3-6-flash - 84. GPT-5 (OpenAI): score 60.4 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gpt-5 - 85. internlm-20b (Unknown): score 60.3 · n/a in / n/a out per 1M · https://benchgecko.ai/model/internlm-20b - 86. DeepSeek V4 Flash 0731 (DeepSeek): score 60.1 · $0.0152 in / $1.28 out per 1M · https://benchgecko.ai/model/deepseek-v4-flash-0731 - 87. Stable Beluga 2 (Unknown): score 60.1 · n/a in / n/a out per 1M · https://benchgecko.ai/model/stable-beluga-2 - 88. Qwen-14B (Alibaba Qwen): score 60.1 · n/a in / n/a out per 1M · https://benchgecko.ai/model/qwen-14b - 89. Qwen2.5 72B Instruct (Alibaba Qwen): score 59.6 · $0.36 in / $0.40 out per 1M · https://benchgecko.ai/model/qwen-2-5-72b-instruct - 90. MiMo-V2-Pro (xiaomi): score 59.6 · $1.00 in / $3.00 out per 1M · https://benchgecko.ai/model/mimo-v2-pro - 91. Claude Opus 4.5 (Anthropic): score 59 · $5.00 in / $25.00 out per 1M · https://benchgecko.ai/model/claude-opus-4-5 - 92. Nemotron-4 15B (Unknown): score 58.9 · n/a in / n/a out per 1M · https://benchgecko.ai/model/nemotron-4-15b - 93. DeepSeek R1 Distill Qwen 32B (DeepSeek): score 58.8 · n/a in / n/a out per 1M · https://benchgecko.ai/model/deepseek-ai-deepseek-r1-distill-qwen-32b - 94. Gemini 2.5 Pro (Google DeepMind): score 58.7 · $1.25 in / $10.00 out per 1M · https://benchgecko.ai/model/gemini-2-5-pro - 95. Grok 4 (xAI): score 58.7 · n/a in / n/a out per 1M · https://benchgecko.ai/model/grok-4 - 96. DeepSeek V4.1 Flash (DeepSeek): score 58.7 · $0.15 in / $0.60 out per 1M · https://benchgecko.ai/model/deepseek-v4-1-flash - 97. GPT-4o (2024-05-13) (OpenAI): score 58.5 · $5.00 in / $15.00 out per 1M · https://benchgecko.ai/model/gpt-4o-2024-05-13 - 98. GPT-5.2 (OpenAI): score 58.4 · $1.75 in / $14.00 out per 1M · https://benchgecko.ai/model/gpt-5-2 - 99. Qwen3.6 Max Preview (Alibaba Qwen): score 58.4 · $1.03 in / $6.16 out per 1M · https://benchgecko.ai/model/qwen3-6-max-preview - 100. Kimi K2 Thinking (moonshotai): score 58.3 · $0.60 in / $2.50 out per 1M · https://benchgecko.ai/model/kimi-k2-thinking ## Same model, different price: widest provider spreads (as of 2026-10-05) - deepseek-v4-1-flash: 28 providers, cheapest Relace $0.0030 in, priciest Fireworks $0.45 in (spread 14900%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v4-1-flash - glm-5-3: 32 providers, cheapest Relace $0.0300 in, priciest Alibaba $2.80 in (spread 9233%) · https://benchgecko.ai/pricing/arbitrage/glm-5-3 - qwen3-8-27b: 17 providers, cheapest Wafer $0.0240 in, priciest Cerebras $0.99 in (spread 4025%) · https://benchgecko.ai/pricing/arbitrage/qwen3-8-27b - deepseek-v4-flash-0731: 26 providers, cheapest Relace $0.0152 in, priciest Phala $0.44 in (spread 2795%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v4-flash-0731 - deepseek-v4-flash: 16 providers, cheapest Relace $0.0300 in, priciest Cloudflare $0.44 in (spread 1367%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v4-flash - deepseek-v3-2: 13 providers, cheapest GMICloud $0.21 in, priciest SambaNova $3.00 in (spread 1337%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v3-2 - gpt-6-astra: 3 providers, cheapest OpenAI $5.00 in, priciest OpenAI $60.00 in (spread 1100%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-astra - gpt-6-astra-pro: 2 providers, cheapest OpenAI $5.00 in, priciest OpenAI $60.00 in (spread 1100%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-astra-pro - gpt-oss-120b: 20 providers, cheapest CoreWeave $0.0300 in, priciest Cerebras $0.35 in (spread 1067%) · https://benchgecko.ai/pricing/arbitrage/gpt-oss-120b - llama-3-1-8b-instruct: 5 providers, cheapest DeepInfra $0.0200 in, priciest CoreWeave $0.22 in (spread 1000%) · https://benchgecko.ai/pricing/arbitrage/llama-3-1-8b-instruct - llama-3-3-70b-instruct: 10 providers, cheapest DeepInfra $0.10 in, priciest Together $1.04 in (spread 940%) · https://benchgecko.ai/pricing/arbitrage/llama-3-3-70b-instruct - deepseek-v4-pro: 16 providers, cheapest Relace $0.21 in, priciest Azure $1.91 in (spread 824%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v4-pro - deepseek-v4-pro-0813: 20 providers, cheapest Relace $0.19 in, priciest Venice $1.65 in (spread 768%) · https://benchgecko.ai/pricing/arbitrage/deepseek-v4-pro-0813 - glm-5-3-flash: 31 providers, cheapest Relace $0.0352 in, priciest Cloudflare $0.30 in (spread 752%) · https://benchgecko.ai/pricing/arbitrage/glm-5-3-flash - gemma-4-31b-it: 12 providers, cheapest DeepInfra $0.0900 in, priciest SiliconFlow $0.75 in (spread 733%) · https://benchgecko.ai/pricing/arbitrage/gemma-4-31b-it - mistral-nemo: 6 providers, cheapest DekaLLM $0.0180 in, priciest Mistral $0.15 in (spread 733%) · https://benchgecko.ai/pricing/arbitrage/mistral-nemo - kimi-k3: 20 providers, cheapest Relace $0.63 in, priciest Fireworks $4.50 in (spread 620%) · https://benchgecko.ai/pricing/arbitrage/kimi-k3 - gpt-5-5: 3 providers, cheapest OpenAI $2.50 in, priciest OpenAI $12.50 in (spread 400%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-5 - mythomax-l2-13b: 3 providers, cheapest Parasail $0.0800 in, priciest DeepInfra $0.40 in (spread 400%) · https://benchgecko.ai/pricing/arbitrage/mythomax-l2-13b - qwen3-6-35b-a3b: 9 providers, cheapest Darkbloom $0.0500 in, priciest CoreWeave $0.25 in (spread 400%) · https://benchgecko.ai/pricing/arbitrage/qwen3-6-35b-a3b - qwen3-coder: 5 providers, cheapest Google $0.22 in, priciest Alibaba $0.97 in (spread 343%) · https://benchgecko.ai/pricing/arbitrage/qwen3-coder - gpt-5-6-sol: 3 providers, cheapest OpenAI $1.00 in, priciest Azure $4.40 in (spread 340%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-sol - gpt-5-6-sol-pro: 2 providers, cheapest OpenAI $1.00 in, priciest Azure $4.40 in (spread 340%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-sol-pro - qwen3-coder-30b-a3b-instruct: 4 providers, cheapest Novita $0.0700 in, priciest Alibaba $0.29 in (spread 318%) · https://benchgecko.ai/pricing/arbitrage/qwen3-coder-30b-a3b-instruct - gpt-oss-20b: 11 providers, cheapest Darkbloom $0.0180 in, priciest Groq $0.0750 in (spread 317%) · https://benchgecko.ai/pricing/arbitrage/gpt-oss-20b - gpt-5-1: 2 providers, cheapest OpenAI $0.63 in, priciest OpenAI $2.50 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-1 - gpt-5-2: 2 providers, cheapest OpenAI $0.88 in, priciest OpenAI $3.50 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-2 - gpt-5-4: 3 providers, cheapest OpenAI $1.25 in, priciest OpenAI $5.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-4 - gpt-5-4-mini: 2 providers, cheapest OpenAI $0.38 in, priciest OpenAI $1.50 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-4-mini - gpt-5-6-luna: 3 providers, cheapest OpenAI $0.10 in, priciest OpenAI $0.40 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-luna - gpt-5-6-luna-pro: 2 providers, cheapest OpenAI $0.10 in, priciest OpenAI $0.40 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-luna-pro - gpt-5-6-terra: 3 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-terra - gpt-5-6-terra-pro: 2 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-5-6-terra-pro - gpt-6-1-sol: 3 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-1-sol - gpt-6-1-sol-pro: 2 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-1-sol-pro - gpt-6-luna: 3 providers, cheapest OpenAI $0.0500 in, priciest OpenAI $0.20 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-luna - gpt-6-luna-pro: 2 providers, cheapest OpenAI $0.0500 in, priciest OpenAI $0.20 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-luna-pro - gpt-6-sol: 3 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-sol - gpt-6-sol-pro: 2 providers, cheapest OpenAI $1.00 in, priciest OpenAI $4.00 in (spread 300%) · https://benchgecko.ai/pricing/arbitrage/gpt-6-sol-pro - qwen3-5-35b-a3b: 7 providers, cheapest Darkbloom $0.0800 in, priciest Venice $0.31 in (spread 291%) · https://benchgecko.ai/pricing/arbitrage/qwen3-5-35b-a3b ## AI companies by valuation (facts as of 2026-10-02) - NVIDIA: market cap $5.64T · as of 2026-10-02, source SEC shares outstanding (2026-08-21) × NVDA close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1045810 · https://benchgecko.ai/economy/company/nvidia - Alphabet (Google): market cap $4.20T · as of 2026-10-02, source SEC common shares outstanding (2026-06-30) × GOOGL close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1652044 · https://benchgecko.ai/economy/company/google - Microsoft: market cap $3.84T · as of 2026-10-02, source SEC shares outstanding (2026-07-23) × MSFT close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=789019 · https://benchgecko.ai/economy/company/microsoft - Amazon (AWS): market cap $2.71T · as of 2026-10-02, source SEC shares outstanding (2026-07-22) × AMZN close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1018724 · https://benchgecko.ai/economy/company/amazon-aws - Meta Platforms: market cap $1.87T · as of 2026-10-02, source SEC diluted weighted shares (2026-06-30) × META close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1326801 · https://benchgecko.ai/economy/company/meta - Broadcom: market cap $1.70T, AI revenue run rate $56.0B · as of 2026-10-02, source SEC shares outstanding (2026-08-28) × AVGO close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1730168 · https://benchgecko.ai/economy/company/broadcom - Micron Technology: market cap $1.21T, AI revenue run rate $30.0B · as of 2026-10-02, source SEC shares outstanding (2026-06-17) × MU close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=723125 · https://benchgecko.ai/economy/company/micron - AMD: market cap $1.03T, AI revenue run rate $28.0B · as of 2026-10-02, source SEC shares outstanding (2026-07-29) × AMD close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=2488 · https://benchgecko.ai/economy/company/amd - Anthropic: valuation $965.0B, AI revenue run rate $14.0B · as of 2026-05-28, source nbcnews.com https://www.nbcnews.com/tech/tech-news/anthropic-secures-965-billion-valuation-after-raising-65-billion-rcna347390 · https://benchgecko.ai/economy/company/anthropic - OpenAI: valuation $840.0B, AI revenue run rate $25.0B · as of 2026-02-01 · https://benchgecko.ai/economy/company/openai - Intel: market cap $601.9B, AI revenue run rate $55.0B · as of 2026-10-02, source SEC shares outstanding (2026-07-17) × INTC close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=50863 · https://benchgecko.ai/economy/company/intel - xAI: valuation $250.0B, AI revenue run rate $500M · as of 2026-02-01 · https://benchgecko.ai/economy/company/xai - Marvell Technology: market cap $238.8B, AI revenue run rate $6.0B · as of 2026-10-02, source SEC shares outstanding (2026-08-21) × MRVL close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=1835632 · https://benchgecko.ai/economy/company/marvell - Qualcomm: market cap $194.1B, AI revenue run rate $42.0B · as of 2026-10-02, source SEC shares outstanding (2026-07-27) × QCOM close https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=804328 · https://benchgecko.ai/economy/company/qualcomm - Databricks: valuation $190.0B, AI revenue run rate $7.0B · as of 2026-08-13, source databricks.com https://www.databricks.com/company/newsroom/press-releases/databricks-grows-80-yoy-surpasses-7b-revenue-run-rate-scales · https://benchgecko.ai/economy/company/databricks - CoreWeave: market cap $49.4B · as of 2026-03-01 · https://benchgecko.ai/economy/company/coreweave - Zhipu AI: market cap $31.3B, AI revenue run rate $53M · as of 2026-03-01 · https://benchgecko.ai/economy/company/zhipu-ai - Crusoe Energy: valuation $30.9B · as of 2026-09-17, source crusoe.ai https://www.crusoe.ai/resources/newsroom/crusoe-announces-series-f-funding · https://benchgecko.ai/economy/company/crusoe-energy - Anysphere (Cursor): valuation $29.3B, AI revenue run rate $1.0B · as of 2025-11-01 · https://benchgecko.ai/economy/company/anysphere - Scale AI: valuation $29.0B, AI revenue run rate $2.0B · as of 2025-06-12, source scale.com https://scale.com/about · https://benchgecko.ai/economy/company/scale-ai - Cerebras Systems: valuation $23.0B, AI revenue run rate $1.0B · as of 2026-02-03, source cerebras.ai https://www.cerebras.ai/press-release/cerebras-systems-raises-usd1-billion-series-h · https://benchgecko.ai/economy/company/cerebras - ElevenLabs: valuation $22.0B, AI revenue run rate $500M · as of 2026-09-30, source elevenlabs.io https://elevenlabs.io/blog/tender-22bn · https://benchgecko.ai/economy/company/elevenlabs - Perplexity AI: valuation $21.2B, AI revenue run rate $200M · as of 2026-01-01 · https://benchgecko.ai/economy/company/perplexity - Moonshot AI: valuation $18.0B · as of 2026-03-01 · https://benchgecko.ai/economy/company/moonshot-ai - Fireworks AI: valuation $17.5B, AI revenue run rate $1.0B · as of 2026-07-16, source cnbc.com https://www.cnbc.com/2026/07/16/fireworks-nvidia-cloud-ai-startup-value.html · https://benchgecko.ai/economy/company/fireworks-ai - Harvey AI: valuation $15.5B, AI revenue run rate $190M · as of 2026-09-09, source harvey.ai https://www.harvey.ai/blog/harvey-raises-dollar550m-at-a-dollar155b-valuation-to-help-legal-teams-own-their-intelligence · https://benchgecko.ai/economy/company/harvey-ai - IREN: market cap $14.9B · as of 2026-04-01 · https://benchgecko.ai/economy/company/iren - Mistral AI: valuation $14.0B, AI revenue run rate $300M · as of 2026-01-01 · https://benchgecko.ai/economy/company/mistral-ai - SambaNova Systems: valuation $11.0B · as of 2026-07-08, source sambanova.ai https://sambanova.ai/press/sambanova-completes-first-close-of-1b-financing-at-11b-valuation · https://benchgecko.ai/economy/company/sambanova - Midjourney: valuation $10.0B, AI revenue run rate $500M · as of 2025-12-01 · https://benchgecko.ai/economy/company/midjourney - Nebius: market cap $9.5B · as of 2026-04-01 · https://benchgecko.ai/economy/company/nebius - Replit: valuation $9.0B, AI revenue run rate $265M · as of 2026-01-01 · https://benchgecko.ai/economy/company/replit - Together AI: valuation $8.3B, AI revenue run rate $300M · as of 2026-07-01, source finance.yahoo.com https://finance.yahoo.com/technology/ai/articles/together-ai-raises-800-million-180132872.html · https://benchgecko.ai/economy/company/together-ai - Applied Digital: market cap $7.4B · as of 2026-04-01 · https://benchgecko.ai/economy/company/applied-digital - Glean: valuation $7.2B, AI revenue run rate $300M · as of 2025-06-01 · https://benchgecko.ai/economy/company/glean - Cohere: valuation $7.0B, AI revenue run rate $150M · as of 2025-09-01 · https://benchgecko.ai/economy/company/cohere - Runway: valuation $5.3B, AI revenue run rate $90M · as of 2026-02-01 · https://benchgecko.ai/economy/company/runway - Baseten: valuation $5.0B · as of 2026-01-01 · https://benchgecko.ai/economy/company/baseten - Lambda: valuation $5.0B, AI revenue run rate $500M · as of 2025-11-01 · https://benchgecko.ai/economy/company/lambda - StepFun: valuation $5.0B · as of 2026-01-01 · https://benchgecko.ai/economy/company/stepfun ## GPU rental prices (per GPU hour, vast.ai marketplace, as of 2026-10-05) - b300: median $9.94, from $8.13 across 3 offers (2026-10-05) · https://benchgecko.ai/hardware/b300 - b200: median $7.81, from $6.25 across 15 offers (2026-10-05) · https://benchgecko.ai/hardware/b200 - h200: median $4.34, from $3.67 across 32 offers (2026-10-05) · https://benchgecko.ai/hardware/h200 - h100: median $2.36, from $1.25 across 30 offers (2026-10-05) · https://benchgecko.ai/hardware/h100 - a100: median $0.80, from $0.40 across 82 offers (2026-10-05) · https://benchgecko.ai/hardware/a100 ## Benchmarks - Aider · Code Editing (coding) · https://benchgecko.ai/benchmark/aider-edit - Aider polyglot (coding) · https://benchgecko.ai/benchmark/aider-polyglot - ANLI (knowledge) · https://benchgecko.ai/benchmark/anli - APEX-Agents (agentic) · https://benchgecko.ai/benchmark/apex-agents - ARC AI2 (knowledge) · https://benchgecko.ai/benchmark/arc-ai2 - ARC-AGI (reasoning) · https://benchgecko.ai/benchmark/arc-agi - ARC-AGI-2 (reasoning) · https://benchgecko.ai/benchmark/arc-agi-2 - Artificial Analysis · Agentic Index (speed) · https://benchgecko.ai/benchmark/aa-agentic-index - Artificial Analysis · Coding Index (speed) · https://benchgecko.ai/benchmark/aa-coding-index - Artificial Analysis · Quality Index (speed) · https://benchgecko.ai/benchmark/aa-quality-index - Artificial Analysis · CritPt (speed) · https://benchgecko.ai/benchmark/aa-critpt - Artificial Analysis · GDPval (speed) · https://benchgecko.ai/benchmark/aa-gdpval - Artificial Analysis · GPQA Diamond (speed) · https://benchgecko.ai/benchmark/aa-gpqa-diamond - Artificial Analysis · Humanity's Last Exam (speed) · https://benchgecko.ai/benchmark/aa-humanitys-last-exam - Artificial Analysis · IFBench (speed) · https://benchgecko.ai/benchmark/aa-ifbench - Artificial Analysis · Long Context Reasoning (speed) · https://benchgecko.ai/benchmark/aa-long-context-reasoning - Artificial Analysis · MMMU Pro (speed) · https://benchgecko.ai/benchmark/aa-mmmu-pro - Artificial Analysis · SciCode (speed) · https://benchgecko.ai/benchmark/aa-scicode - Artificial Analysis · tau2-Bench Telecom (speed) · https://benchgecko.ai/benchmark/aa-tau2-bench - Artificial Analysis · Terminal-Bench Hard (speed) · https://benchgecko.ai/benchmark/aa-terminal-bench-hard - AudioMultiChallenge (knowledge) · https://benchgecko.ai/benchmark/seal-audiomultichallenge - AudioMultiChallenge · Audio Output (knowledge) · https://benchgecko.ai/benchmark/seal-audiomultichallenge-audio-output - AudioMultiChallenge · Text Output (knowledge) · https://benchgecko.ai/benchmark/seal-audiomultichallenge-text-output - Balrog (knowledge) · https://benchgecko.ai/benchmark/balrog - BBH (reasoning) · https://benchgecko.ai/benchmark/bbh - BBH (HuggingFace) (general) · https://benchgecko.ai/benchmark/hf-bbh - C-Eval (knowledge) · https://benchgecko.ai/benchmark/c-eval - CadEval (coding) · https://benchgecko.ai/benchmark/cadeval - CharXiv Reasoning (reasoning) · https://benchgecko.ai/benchmark/charxiv-reasoning - CharXiv Reasoning (with tools) (reasoning) · https://benchgecko.ai/benchmark/charxiv-reasoning-tools - Chatbot Arena Elo · Coding (arena) · https://benchgecko.ai/benchmark/arena-elo-coding - Chatbot Arena Elo · Overall (arena) · https://benchgecko.ai/benchmark/arena-elo-overall - Chess Puzzles (knowledge) · https://benchgecko.ai/benchmark/chess-puzzles - Cl Bench (general) · https://benchgecko.ai/benchmark/cl-bench - Cl Bench Life (general) · https://benchgecko.ai/benchmark/cl-bench-life - CMMLU (knowledge) · https://benchgecko.ai/benchmark/cmmlu - CSQA2 (knowledge) · https://benchgecko.ai/benchmark/csqa2 - Cybench (coding) · https://benchgecko.ai/benchmark/cybench - DeepResearch Bench (knowledge) · https://benchgecko.ai/benchmark/deepresearch-bench - Deepswe (coding) · https://benchgecko.ai/benchmark/deepswe - Dtbench (general) · https://benchgecko.ai/benchmark/dtbench - Ebr Bench (general) · https://benchgecko.ai/benchmark/ebr-bench - EnigmaEval (knowledge) · https://benchgecko.ai/benchmark/seal-enigmaeval - Exploitbench (general) · https://benchgecko.ai/benchmark/exploitbench - Fiction.LiveBench (knowledge) · https://benchgecko.ai/benchmark/fiction-livebench - Fortress (safety) · https://benchgecko.ai/benchmark/seal-fortress - Frontiercode (coding) · https://benchgecko.ai/benchmark/frontiercode - FrontierMath-2025-02-28-Private (math) · https://benchgecko.ai/benchmark/frontiermath-2025-02-28-private - FrontierMath-Tier-4-2025-07-01-Private (math) · https://benchgecko.ai/benchmark/frontiermath-tier-4-2025-07-01-private - FrontierMath-Tier-4-v2-Private (math) · https://benchgecko.ai/benchmark/frontiermath-tier-4-v2-private - FrontierMath-Tiers-1-3-v2-Private (math) · https://benchgecko.ai/benchmark/frontiermath-tiers-1-3-v2-private - Frontierswe (coding) · https://benchgecko.ai/benchmark/frontierswe - Furniture Assembly (general) · https://benchgecko.ai/benchmark/furniture-assembly - Gdpval (general) · https://benchgecko.ai/benchmark/gdpval - GeoBench (knowledge) · https://benchgecko.ai/benchmark/geobench - GPQA (knowledge) · https://benchgecko.ai/benchmark/hf-gpqa - GPQA diamond (knowledge) · https://benchgecko.ai/benchmark/gpqa-diamond - GraphWalks BFS 256K-1M (reasoning) · https://benchgecko.ai/benchmark/graphwalks-bfs-256k - GSM8K (math) · https://benchgecko.ai/benchmark/gsm8k - GSO-Bench (coding) · https://benchgecko.ai/benchmark/gso-bench - HellaSwag (knowledge) · https://benchgecko.ai/benchmark/hellaswag - HELM · GPQA (knowledge) · https://benchgecko.ai/benchmark/helm-gpqa - HELM · IFEval (language) · https://benchgecko.ai/benchmark/helm-ifeval - HELM · MMLU-Pro (knowledge) · https://benchgecko.ai/benchmark/helm-mmlu-pro - HELM · Omni-MATH (math) · https://benchgecko.ai/benchmark/helm-omni-math - HELM · WildBench (reasoning) · https://benchgecko.ai/benchmark/helm-wildbench - HLE (knowledge) · https://benchgecko.ai/benchmark/hle - HLE (with tools) (reasoning) · https://benchgecko.ai/benchmark/hle-tools - Humanity's Last Exam (knowledge) · https://benchgecko.ai/benchmark/seal-humanitys-last-exam - Humanity's Last Exam (Text Only) (knowledge) · https://benchgecko.ai/benchmark/seal-humanitys-last-exam-text - IFEval (language) · https://benchgecko.ai/benchmark/hf-ifeval - JCommonsenseQA (language) · https://benchgecko.ai/benchmark/jp-jcommonsenseqa - JHumanEval (language) · https://benchgecko.ai/benchmark/jp-jhumaneval - JMMLU (language) · https://benchgecko.ai/benchmark/jp-jmmlu - JNLI (language) · https://benchgecko.ai/benchmark/jp-jnli - JSQuAD (language) · https://benchgecko.ai/benchmark/jp-jsquad - LAMBADA (knowledge) · https://benchgecko.ai/benchmark/lambada - Lech Mazur Writing (knowledge) · https://benchgecko.ai/benchmark/lech-mazur-writing - LiveBench · Agentic Coding (coding) · https://benchgecko.ai/benchmark/livebench-agentic-coding - LiveBench · Coding (coding) · https://benchgecko.ai/benchmark/livebench-coding - LiveBench · Data Analysis (reasoning) · https://benchgecko.ai/benchmark/livebench-data-analysis - LiveBench · If (language) · https://benchgecko.ai/benchmark/livebench-if - LiveBench · Language (language) · https://benchgecko.ai/benchmark/livebench-language - LiveBench · Mathematics (math) · https://benchgecko.ai/benchmark/livebench-mathematics - LiveBench · Overall (knowledge) · https://benchgecko.ai/benchmark/livebench-overall - LiveBench · Reasoning (reasoning) · https://benchgecko.ai/benchmark/livebench-reasoning - LLM-JP · Overall (language) · https://benchgecko.ai/benchmark/jp-overall - Lmca (general) · https://benchgecko.ai/benchmark/lmca - MASK (safety) · https://benchgecko.ai/benchmark/seal-mask - MATH level 5 (math) · https://benchgecko.ai/benchmark/math-level-5 - MATH Level 5 (math) · https://benchgecko.ai/benchmark/hf-math-lvl5 - MCP Atlas (agentic) · https://benchgecko.ai/benchmark/seal-mcp-atlas - Metr Time Horizons (general) · https://benchgecko.ai/benchmark/metr-time-horizons - Mirrorcode (coding) · https://benchgecko.ai/benchmark/mirrorcode - MMLU (knowledge) · https://benchgecko.ai/benchmark/mmlu - MMLU-PRO (knowledge) · https://benchgecko.ai/benchmark/hf-mmlu-pro - MMMLU (knowledge) · https://benchgecko.ai/benchmark/mmmlu - MMMLU · Arabic (language) · https://benchgecko.ai/benchmark/mmmlu-ar - MMMLU · Bengali (language) · https://benchgecko.ai/benchmark/mmmlu-bn - MMMLU · Chinese (language) · https://benchgecko.ai/benchmark/mmmlu-zh - MMMLU · French (language) · https://benchgecko.ai/benchmark/mmmlu-fr - MMMLU · German (language) · https://benchgecko.ai/benchmark/mmmlu-de - MMMLU · Hindi (language) · https://benchgecko.ai/benchmark/mmmlu-hi - MMMLU · Indonesian (language) · https://benchgecko.ai/benchmark/mmmlu-id - MMMLU · Italian (language) · https://benchgecko.ai/benchmark/mmmlu-it - MMMLU · Japanese (language) · https://benchgecko.ai/benchmark/mmmlu-ja - MMMLU · Korean (language) · https://benchgecko.ai/benchmark/mmmlu-ko - MMMLU · Portuguese (language) · https://benchgecko.ai/benchmark/mmmlu-pt - MMMLU · Spanish (language) · https://benchgecko.ai/benchmark/mmmlu-es - MMMLU · Swahili (language) · https://benchgecko.ai/benchmark/mmmlu-sw - MMMLU · Yoruba (language) · https://benchgecko.ai/benchmark/mmmlu-yo - MultiChallenge (knowledge) · https://benchgecko.ai/benchmark/seal-multichallenge - MultiNRC (knowledge) · https://benchgecko.ai/benchmark/seal-multinrc - MUSR (reasoning) · https://benchgecko.ai/benchmark/hf-musr - Mystery Game Puzzles (general) · https://benchgecko.ai/benchmark/mystery-game-puzzles - OpenBookQA (knowledge) · https://benchgecko.ai/benchmark/openbookqa - OpenCompass · AIME2025 (math) · https://benchgecko.ai/benchmark/oc-aime2025 - OpenCompass · GPQA-Diamond (knowledge) · https://benchgecko.ai/benchmark/oc-gpqa-diamond - OpenCompass · HLE (knowledge) · https://benchgecko.ai/benchmark/oc-hle - OpenCompass · IFEval (language) · https://benchgecko.ai/benchmark/oc-ifeval - OpenCompass · LiveCodeBenchV6 (coding) · https://benchgecko.ai/benchmark/oc-livecodebenchv6 - OpenCompass · MMLU-Pro (knowledge) · https://benchgecko.ai/benchmark/oc-mmlu-pro - OSWorld (agentic) · https://benchgecko.ai/benchmark/osworld - Osworld 2 0 (general) · https://benchgecko.ai/benchmark/osworld-2-0 - OTIS Mock AIME 2024-2025 (math) · https://benchgecko.ai/benchmark/otis-mock-aime-2024-2025 - PIQA (knowledge) · https://benchgecko.ai/benchmark/piqa - PostTrainBench (knowledge) · https://benchgecko.ai/benchmark/posttrainbench - Professional Reasoning · Finance (knowledge) · https://benchgecko.ai/benchmark/seal-pro-reasoning-finance - Professional Reasoning · Legal (knowledge) · https://benchgecko.ai/benchmark/seal-pro-reasoning-legal - Proofbench (general) · https://benchgecko.ai/benchmark/proofbench - PropensityBench (safety) · https://benchgecko.ai/benchmark/seal-propensitybench - Remote Labor Index (general) · https://benchgecko.ai/benchmark/remote-labor-index - Remote Labor Index (RLI) (agentic) · https://benchgecko.ai/benchmark/seal-remote-labor-index - ScienceQA (knowledge) · https://benchgecko.ai/benchmark/scienceqa - SciPredict (knowledge) · https://benchgecko.ai/benchmark/seal-scipredict - SimpleBench (reasoning) · https://benchgecko.ai/benchmark/simplebench - SimpleQA Verified (knowledge) · https://benchgecko.ai/benchmark/simpleqa-verified - Surface Evolver Bench (general) · https://benchgecko.ai/benchmark/surface-evolver-bench - SWE Atlas · Codebase QnA (agentic) · https://benchgecko.ai/benchmark/seal-swe-atlas-codebase-qna - SWE Atlas · Test Writing (agentic) · https://benchgecko.ai/benchmark/seal-swe-atlas-test-writing - SWE-bench Multilingual (coding) · https://benchgecko.ai/benchmark/swe-bench-multilingual - SWE-bench Multimodal (coding) · https://benchgecko.ai/benchmark/swe-bench-multimodal - SWE-bench Pro (coding) · https://benchgecko.ai/benchmark/swe-bench-pro - SWE-Bench Pro (Private) (agentic) · https://benchgecko.ai/benchmark/seal-swe-bench-pro-private - SWE-Bench Pro (Public) (agentic) · https://benchgecko.ai/benchmark/seal-swe-bench-pro-public - SWE-Bench verified (coding) · https://benchgecko.ai/benchmark/swe-bench-verified - SWE-Bench Verified (Bash Only) (coding) · https://benchgecko.ai/benchmark/swe-bench-verified-bash-only - Terminal Bench (coding) · https://benchgecko.ai/benchmark/terminal-bench - The Agent Company (agentic) · https://benchgecko.ai/benchmark/the-agent-company - TriviaQA (knowledge) · https://benchgecko.ai/benchmark/triviaqa - TutorBench (knowledge) · https://benchgecko.ai/benchmark/seal-tutorbench - USAMO (math) · https://benchgecko.ai/benchmark/usamo - VideoMME (multimodal) · https://benchgecko.ai/benchmark/videomme - VISTA (knowledge) · https://benchgecko.ai/benchmark/seal-vista - VisualToolBench (VTB) (knowledge) · https://benchgecko.ai/benchmark/seal-visual-tool-bench - VPCT (knowledge) · https://benchgecko.ai/benchmark/vpct - WeirdML (coding) · https://benchgecko.ai/benchmark/weirdml - Winogrande (knowledge) · https://benchgecko.ai/benchmark/winogrande ## AI accelerators - B200 by NVIDIA (Blackwell) · https://benchgecko.ai/hardware/b200 - GB200 by NVIDIA (Grace Blackwell) · https://benchgecko.ai/hardware/gb200 - TPU v7 by Google (Ironwood) · https://benchgecko.ai/hardware/tpu-v7 - Trainium2 by AWS (Trainium2) · https://benchgecko.ai/hardware/trainium2 - GB300 by NVIDIA (Grace Blackwell Ultra) · https://benchgecko.ai/hardware/gb300 - MI355X by AMD (Altair) · https://benchgecko.ai/hardware/mi355x - Ascend 910C by Huawei (Ascend 910C) · https://benchgecko.ai/hardware/ascend-910c - TPU v6e by Google (Trillium) · https://benchgecko.ai/hardware/tpu-v6e - B300 by NVIDIA (Blackwell Ultra) · https://benchgecko.ai/hardware/b300 - MI325X by AMD (Aqua Vanjaram+) · https://benchgecko.ai/hardware/mi325x - MI300X by AMD (Aqua Vanjaram) · https://benchgecko.ai/hardware/mi300x - WSE-3 by Cerebras (WSE-3) · https://benchgecko.ai/hardware/wse-3 - Gaudi 3 by Intel (Gaudi 3) · https://benchgecko.ai/hardware/gaudi-3 - H100 by NVIDIA (Hopper) · https://benchgecko.ai/hardware/h100 - H200 by NVIDIA (Hopper) · https://benchgecko.ai/hardware/h200 - MTIA v2 by Meta (Artemis) · https://benchgecko.ai/hardware/mtia-v2 - TPU v5p by Google (Viperfish) · https://benchgecko.ai/hardware/tpu-v5p - Maia 100 by Microsoft (Maia 100) · https://benchgecko.ai/hardware/maia-100 - LPU by Groq (GroqChip 1) · https://benchgecko.ai/hardware/groq-lpu - A100 by NVIDIA (Ampere) · https://benchgecko.ai/hardware/a100 ## AI glossary ### Learning path · The AI Bubble Explained Seven terms that decode whether AI is overpriced, fairly priced, or criminally underpriced. Read in order. - https://benchgecko.ai/learn/glossary/bubble-index · Start with the aggregate gauge: how stretched are AI valuations? - https://benchgecko.ai/learn/glossary/ps-ratio · Price to sales: the ratio behind most AI valuation debates. - https://benchgecko.ai/learn/glossary/arr · Annual recurring revenue, the number AI companies report most. - https://benchgecko.ai/learn/glossary/burn-rate · How fast labs spend the money they raise. - https://benchgecko.ai/learn/glossary/capex · The infrastructure spending cycle behind the boom. - https://benchgecko.ai/learn/glossary/moe · Why the cost of a token keeps falling. - https://benchgecko.ai/learn/glossary/reasoning-model · The premium tier, and why it matters for pricing. ### Learning path · Pick an AI Model Six terms to go from "I need an AI" to "here is the cheapest model that meets my spec." - https://benchgecko.ai/learn/glossary/mmlu · The baseline knowledge benchmark everyone cites. - https://benchgecko.ai/learn/glossary/swe-bench · If your workload is code, this is the one to care about. - https://benchgecko.ai/learn/glossary/context-window · How much input the model can hold at once. - https://benchgecko.ai/learn/glossary/input-tokens · Price per million input tokens, the biggest line on most bills. - https://benchgecko.ai/learn/glossary/throughput · Speed: tokens per second, which shapes the user experience. - https://benchgecko.ai/learn/glossary/function-calling · How a model connects to your code and tools in production. ### Learning path · From Sand to Model The AI supply chain in 7 terms · foundry, memory, chip, system, training, inference, pricing. - https://benchgecko.ai/learn/glossary/tsmc · Where most AI chips are manufactured. - https://benchgecko.ai/learn/glossary/hbm3e · High bandwidth memory, the throughput backbone. - https://benchgecko.ai/learn/glossary/gpu · Where the silicon becomes an accelerator. - https://benchgecko.ai/learn/glossary/nvidia-dgx-gb200-nvl72 · Many chips wired into one rack-scale system. - https://benchgecko.ai/learn/glossary/pretraining · What these systems are used for first. - https://benchgecko.ai/learn/glossary/inference · What they are used for every day after that. - https://benchgecko.ai/learn/glossary/input-tokens · How model compute finally becomes a price. ### Learning path · What is GenAI · for normies Zero jargon. Six terms to sound smart at a dinner party. - https://benchgecko.ai/learn/glossary/tokens · The thing AI charges for. - https://benchgecko.ai/learn/glossary/transformer · The idea behind every modern AI. - https://benchgecko.ai/learn/glossary/pretraining · How an AI learns, in plain English. - https://benchgecko.ai/learn/glossary/inference · What happens every time you use AI. - https://benchgecko.ai/learn/glossary/prompt-engineering · Why typing matters. - https://benchgecko.ai/learn/glossary/agent · AI that takes actions, not just answers. ### Learning path · The Reasoning Era How reasoning models changed what an answer costs. - https://benchgecko.ai/learn/glossary/chain-of-thought · The prompting trick that became a training technique. - https://benchgecko.ai/learn/glossary/reasoning-model · How models that think before answering work. - https://benchgecko.ai/learn/glossary/test-time-compute · Spending more compute at answer time. - https://benchgecko.ai/learn/glossary/gpqa-diamond · A hard science benchmark reasoning models are judged on. - https://benchgecko.ai/learn/glossary/arc-agi · A puzzle benchmark built to resist memorization. - https://benchgecko.ai/learn/glossary/reasoning-token-billing · Why reasoning answers cost more. ### Learning path · Why AI is Expensive Six terms for the angry question at the VC meeting. - https://benchgecko.ai/learn/glossary/gpu · Supply-constrained and margin-rich. - https://benchgecko.ai/learn/glossary/hbm3e · The memory bottleneck. - https://benchgecko.ai/learn/glossary/fp16 · How compute capacity gets measured. - https://benchgecko.ai/learn/glossary/pretraining · Why frontier training runs are so costly. - https://benchgecko.ai/learn/glossary/output-tokens · Why serving costs never go to zero. - https://benchgecko.ai/learn/glossary/capex · The spending that all this requires. ### Benchmarks #### SWE-bench URL: https://benchgecko.ai/learn/glossary/swe-bench Text reviewed 2026-10-05 TL;DR: A benchmark where models attempt real GitHub issues · judged by whether their patch passes the project's test suite. Each SWE-bench task is a real GitHub issue paired with the exact repo state before the merge. The model must output a patch that makes the hidden test suite pass without regressions. Popular variants: SWE-bench Verified (500 hand-checked tasks, the canonical leaderboard), SWE-bench Lite (300 easier tasks), Multimodal (screenshots), Multilingual (non-Python repos). The leaderboard separates agent-style solutions (tool use + multi-turn) from single-shot patch generation. See live data: https://benchgecko.ai/benchmark/swe-bench-verified #### HumanEval URL: https://benchgecko.ai/learn/glossary/humaneval Text reviewed 2026-10-05 TL;DR: A benchmark where the model writes a Python function from its docstring, scored by whether it passes hidden unit tests. HumanEval's 164 problems span string manipulation, list operations, math, and simple algorithms. The test suite is held out, not shown to the model. Pass@1 is the hardest metric · one shot, must be correct. Pass@10 and pass@100 allow the model multiple attempts. HumanEval+ extends the suite with additional adversarial tests. Real-world coding capability is better measured by SWE-bench, Aider polyglot, or LiveBench coding subset. See live data: https://benchgecko.ai/benchmark/humaneval #### LiveBench URL: https://benchgecko.ai/learn/glossary/livebench Text reviewed 2026-10-05 TL;DR: A benchmark that refreshes its tasks monthly so models can't memorize them · tests reasoning, coding, math, data, and instruction following. LiveBench ships monthly new-task releases. The leaderboard shows both "recent" scores (last 3 months) and "all-time" scores. Categories weight equally: reasoning (logic puzzles, planning), coding (novel problems beyond HumanEval), mathematics (AMC-style), data analysis (table-based questions), instruction following (strict format compliance), and language (complex writing). The "all-time" leaderboard ages out contaminated scores as tasks are deprecated. See live data: https://benchgecko.ai/benchmark/livebench #### Chatbot Arena URL: https://benchgecko.ai/learn/glossary/chatbot-arena TL;DR: A blind pairwise AI model comparison where humans vote · outputs are anonymous and Elo ratings produce the leaderboard. Chatbot Arena (LMSYS, UC Berkeley) has collected 15M+ votes since 2023. The rating is Bradley-Terry Elo updated in real time. Models are presented anonymously as "Model A" vs "Model B"; users choose A, B, tie, or both bad. Category filters (coding, reasoning, hard prompts, math) slice the leaderboard by query type. The overall leaderboard is dominated by general-purpose chat quality; specialized categories rank differently. Methodological concerns: voter self-selection bias, short queries dominating, style-over-substance preferences. Despite critiques, Arena is the closest thing to a real-world preference measure in 2026. See live data: https://benchgecko.ai/benchmark/chatbot-arena #### MMLU Pro URL: https://benchgecko.ai/learn/glossary/mmlu-pro Text reviewed 2026-10-05 TL;DR: A harder, more discriminating version of MMLU with 10 answer choices instead of 4 and reasoning-heavy questions. MMLU Pro filters trivial and ambiguous questions from the original MMLU and adds new reasoning-heavy items. Each question has 10 options to reduce lucky-guess inflation from 25% to 10% baseline. Subject coverage: Math, Physics, Chemistry, Biology, Computer Science, Engineering, Economics, Business, Law, Psychology, Philosophy, History, Health, and Other. The "Math" subset is the hardest · top models score below 70%. Pro reports significantly larger gaps between frontier and mid-tier models than plain MMLU, making it the better discriminator in 2026. See live data: https://benchgecko.ai/benchmark/mmlu-pro #### ARC AI2 URL: https://benchgecko.ai/learn/glossary/arc-ai2 Built from BenchGecko data as of 2026-10-02 TL;DR: ARC AI2 is a knowledge benchmark tracked on BenchGecko. AI2 Reasoning Challenge. Grade-school science questions requiring multi-step reasoning. Easy and Challenge sets test different difficulty levels. Scores are reported (%, maximum 100); higher is better. BenchGecko collects ARC AI2 scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/arc-ai2 #### BBH URL: https://benchgecko.ai/learn/glossary/bbh Built from BenchGecko data as of 2026-06-22 TL;DR: BBH is a reasoning benchmark tracked on BenchGecko. BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance. Scores are reported (%, maximum 100); higher is better. BenchGecko collects BBH scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/bbh #### GSM8K URL: https://benchgecko.ai/learn/glossary/gsm8k Built from BenchGecko data as of 2026-06-22 TL;DR: GSM8K is a math benchmark tracked on BenchGecko. Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark. Scores are reported (%, maximum 100); higher is better. BenchGecko collects GSM8K scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/gsm8k #### HellaSwag URL: https://benchgecko.ai/learn/glossary/hellaswag Built from BenchGecko data as of 2026-06-22 TL;DR: HellaSwag is a knowledge benchmark tracked on BenchGecko. Sentence completion requiring commonsense reasoning about physical and social situations. Tests real-world understanding. Scores are reported (%, maximum 100); higher is better. BenchGecko collects HellaSwag scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/hellaswag #### LAMBADA URL: https://benchgecko.ai/learn/glossary/lambada Built from BenchGecko data as of 2026-06-22 TL;DR: LAMBADA is a knowledge benchmark tracked on BenchGecko. Language modeling benchmark testing ability to predict the last word of passages requiring long-range context understanding. Scores are reported (%, maximum 100); higher is better. BenchGecko collects LAMBADA scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/lambada #### MMLU URL: https://benchgecko.ai/learn/glossary/mmlu Built from BenchGecko data as of 2026-10-02 TL;DR: MMLU is a knowledge benchmark tracked on BenchGecko. Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark. Scores are reported (%, maximum 100); higher is better. BenchGecko collects MMLU scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/mmlu #### GPQA diamond URL: https://benchgecko.ai/learn/glossary/gpqa-diamond Built from BenchGecko data as of 2026-10-02 TL;DR: GPQA diamond is a knowledge benchmark tracked on BenchGecko. Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding. Scores are reported (%, maximum 100); higher is better. BenchGecko collects GPQA diamond scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/gpqa-diamond #### MATH level 5 URL: https://benchgecko.ai/learn/glossary/math-level-5 Built from BenchGecko data as of 2026-10-02 TL;DR: MATH level 5 is a math benchmark tracked on BenchGecko. Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving. Scores are reported (%, maximum 100); higher is better. BenchGecko collects MATH level 5 scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/math-level-5 #### OTIS Mock AIME 2024-2025 URL: https://benchgecko.ai/learn/glossary/otis-mock-aime-2024-2025 Built from BenchGecko data as of 2026-10-02 TL;DR: OTIS Mock AIME 2024-2025 is a math benchmark tracked on BenchGecko. Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance. Scores are reported (%, maximum 100); higher is better. BenchGecko collects OTIS Mock AIME 2024-2025 scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/otis-mock-aime-2024-2025 #### WeirdML URL: https://benchgecko.ai/learn/glossary/weirdml Built from BenchGecko data as of 2026-10-02 TL;DR: WeirdML is a coding benchmark tracked on BenchGecko. Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems. Scores are reported (%, maximum 100); higher is better. BenchGecko collects WeirdML scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/weirdml #### Winogrande URL: https://benchgecko.ai/learn/glossary/winogrande Built from BenchGecko data as of 2026-06-22 TL;DR: Winogrande is a knowledge benchmark tracked on BenchGecko. Commonsense coreference resolution. Tests understanding of pronoun references in ambiguous sentences. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Winogrande scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/winogrande #### SimpleBench URL: https://benchgecko.ai/learn/glossary/simplebench Built from BenchGecko data as of 2026-10-02 TL;DR: SimpleBench is a reasoning benchmark tracked on BenchGecko. Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps. Scores are reported (%, maximum 100); higher is better. BenchGecko collects SimpleBench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/simplebench #### FrontierMath-2025-02-28-Private URL: https://benchgecko.ai/learn/glossary/frontiermath-2025-02-28-private Built from BenchGecko data as of 2026-06-22 TL;DR: FrontierMath-2025-02-28-Private is a math benchmark tracked on BenchGecko. Original research-level math problems created by professional mathematicians. Problems are unpublished and cannot be memorized. Scores are reported (%, maximum 100); higher is better. BenchGecko collects FrontierMath-2025-02-28-Private scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/frontiermath-2025-02-28-private #### Lech Mazur Writing URL: https://benchgecko.ai/learn/glossary/lech-mazur-writing Built from BenchGecko data as of 2026-10-02 TL;DR: Lech Mazur Writing is a knowledge benchmark tracked on BenchGecko. Writing quality evaluation by Lech Mazur. Tests prose quality, coherence, and stylistic ability. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Lech Mazur Writing scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/lech-mazur-writing #### SimpleQA Verified URL: https://benchgecko.ai/learn/glossary/simpleqa-verified Built from BenchGecko data as of 2026-10-02 TL;DR: SimpleQA Verified is a knowledge benchmark tracked on BenchGecko. Simple factual questions with verified correct answers. Tests accuracy of basic knowledge retrieval. Low scores indicate hallucination. Scores are reported (%, maximum 100); higher is better. BenchGecko collects SimpleQA Verified scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/simpleqa-verified #### Aider polyglot URL: https://benchgecko.ai/learn/glossary/aider-polyglot Built from BenchGecko data as of 2026-10-02 TL;DR: Aider polyglot is a coding benchmark tracked on BenchGecko. Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Aider polyglot scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/aider-polyglot #### ARC-AGI URL: https://benchgecko.ai/learn/glossary/arc-agi Built from BenchGecko data as of 2026-10-02 TL;DR: ARC-AGI is a reasoning benchmark tracked on BenchGecko. Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence. Scores are reported (%, maximum 100); higher is better. BenchGecko collects ARC-AGI scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/arc-agi #### Fiction.LiveBench URL: https://benchgecko.ai/learn/glossary/fiction-livebench Built from BenchGecko data as of 2026-10-02 TL;DR: Fiction.LiveBench is a knowledge benchmark tracked on BenchGecko. LiveBench fiction analysis. Tests literary comprehension and creative text understanding. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Fiction.LiveBench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/fiction-livebench #### ARC-AGI-2 URL: https://benchgecko.ai/learn/glossary/arc-agi-2 Built from BenchGecko data as of 2026-10-02 TL;DR: ARC-AGI-2 is a reasoning benchmark tracked on BenchGecko. ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data. Scores are reported (%, maximum 100); higher is better. BenchGecko collects ARC-AGI-2 scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/arc-agi-2 #### GSO-Bench URL: https://benchgecko.ai/learn/glossary/gso-bench Built from BenchGecko data as of 2026-10-02 TL;DR: GSO-Bench is a coding benchmark tracked on BenchGecko. GitHub Star Optimization benchmark. Evaluates ability to understand and improve open-source projects to increase quality and adoption. Scores are reported (%, maximum 100); higher is better. BenchGecko collects GSO-Bench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/gso-bench #### Terminal Bench URL: https://benchgecko.ai/learn/glossary/terminal-bench Built from BenchGecko data as of 2026-10-02 TL;DR: Terminal Bench is a coding benchmark tracked on BenchGecko. Complex terminal-based engineering tasks. Models must use command-line tools, navigate filesystems, and debug systems through shell interaction. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Terminal Bench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/terminal-bench #### FrontierMath-Tier-4-2025-07-01-Private URL: https://benchgecko.ai/learn/glossary/frontiermath-tier-4-2025-07-01-private Built from BenchGecko data as of 2026-06-22 TL;DR: FrontierMath-Tier-4-2025-07-01-Private is a math benchmark tracked on BenchGecko. Hardest tier of FrontierMath. Problems at the frontier of human mathematical ability, many unsolved by most mathematicians. Scores are reported (%, maximum 100); higher is better. BenchGecko collects FrontierMath-Tier-4-2025-07-01-Private scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/frontiermath-tier-4-2025-07-01-private #### Chess Puzzles URL: https://benchgecko.ai/learn/glossary/chess-puzzles Built from BenchGecko data as of 2026-10-02 TL;DR: Chess Puzzles is a knowledge benchmark tracked on BenchGecko. Tactical chess puzzles testing pattern recognition and multi-move calculation. Measures strategic reasoning ability. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Chess Puzzles scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/chess-puzzles #### APEX-Agents URL: https://benchgecko.ai/learn/glossary/apex-agents Built from BenchGecko data as of 2026-10-02 TL;DR: APEX-Agents is an agentic benchmark tracked on BenchGecko. Agent performance evaluation testing multi-step tool use, planning, and execution in realistic environments. Scores are reported (%, maximum 100); higher is better. BenchGecko collects APEX-Agents scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/apex-agents #### PostTrainBench URL: https://benchgecko.ai/learn/glossary/posttrainbench Built from BenchGecko data as of 2026-10-02 TL;DR: PostTrainBench is a knowledge benchmark tracked on BenchGecko. Evaluates post-training behaviors including instruction following, safety, and helpfulness balance. Scores are reported (%, maximum 100); higher is better. BenchGecko collects PostTrainBench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/posttrainbench #### SWE-Bench verified URL: https://benchgecko.ai/learn/glossary/swe-bench-verified Built from BenchGecko data as of 2026-10-02 TL;DR: SWE-Bench verified is a coding benchmark tracked on BenchGecko. Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench. Scores are reported (%, maximum 100); higher is better. BenchGecko collects SWE-Bench verified scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/swe-bench-verified #### OSWorld URL: https://benchgecko.ai/learn/glossary/osworld Built from BenchGecko data as of 2026-06-22 TL;DR: OSWorld is an agentic benchmark tracked on BenchGecko. Operating System World. Tests ability to complete real computer tasks across Windows, macOS, and Linux environments. Scores are reported (%, maximum 100); higher is better. BenchGecko collects OSWorld scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/osworld #### HLE URL: https://benchgecko.ai/learn/glossary/hle Built from BenchGecko data as of 2026-10-02 TL;DR: HLE is a knowledge benchmark tracked on BenchGecko. Humanitys Last Exam. Expert-level questions spanning all academic disciplines, designed to be the hardest knowledge test for AI. Scores are reported (%, maximum 100); higher is better. BenchGecko collects HLE scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/hle #### TriviaQA URL: https://benchgecko.ai/learn/glossary/triviaqa Built from BenchGecko data as of 2026-10-02 TL;DR: TriviaQA is a knowledge benchmark tracked on BenchGecko. Trivia questions sourced from trivia enthusiasts and quiz websites. Tests breadth of general knowledge. Scores are reported (%, maximum 100); higher is better. BenchGecko collects TriviaQA scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/triviaqa #### ScienceQA URL: https://benchgecko.ai/learn/glossary/scienceqa Built from BenchGecko data as of 2026-04-09 TL;DR: ScienceQA is a knowledge benchmark tracked on BenchGecko. Science questions with multimodal context including diagrams and charts from K-12 curriculum. Scores are reported (%, maximum 100); higher is better. BenchGecko collects ScienceQA scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/scienceqa #### PIQA URL: https://benchgecko.ai/learn/glossary/piqa Built from BenchGecko data as of 2026-06-22 TL;DR: PIQA is a knowledge benchmark tracked on BenchGecko. Physical Intuition QA. Tests understanding of everyday physical interactions and commonsense physics. Scores are reported (%, maximum 100); higher is better. BenchGecko collects PIQA scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/piqa #### OpenBookQA URL: https://benchgecko.ai/learn/glossary/openbookqa Built from BenchGecko data as of 2026-10-02 TL;DR: OpenBookQA is a knowledge benchmark tracked on BenchGecko. Elementary science questions with access to a small book of core science facts. Tests reasoning beyond memorization. Scores are reported (%, maximum 100); higher is better. BenchGecko collects OpenBookQA scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/openbookqa #### CadEval URL: https://benchgecko.ai/learn/glossary/cadeval Built from BenchGecko data as of 2026-06-22 TL;DR: CadEval is a coding benchmark tracked on BenchGecko. Computer-aided design evaluation. Tests understanding of CAD concepts, 3D modeling, and engineering design principles. Scores are reported (%, maximum 100); higher is better. BenchGecko collects CadEval scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/cadeval #### Balrog URL: https://benchgecko.ai/learn/glossary/balrog Built from BenchGecko data as of 2026-10-02 TL;DR: Balrog is a knowledge benchmark tracked on BenchGecko. Broad Assessment of Language and Reasoning Over Games. Tests strategic and logical reasoning through game scenarios. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Balrog scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/balrog #### GeoBench URL: https://benchgecko.ai/learn/glossary/geobench Built from BenchGecko data as of 2026-06-22 TL;DR: GeoBench is a knowledge benchmark tracked on BenchGecko. Geography benchmark testing knowledge of world geography, landmarks, borders, and geopolitical facts. Scores are reported (%, maximum 100); higher is better. BenchGecko collects GeoBench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/geobench #### CSQA2 URL: https://benchgecko.ai/learn/glossary/csqa2 Built from BenchGecko data as of 2026-04-09 TL;DR: CSQA2 is a knowledge benchmark tracked on BenchGecko. CommonsenseQA 2. Tests commonsense reasoning through yes/no questions about everyday situations. Scores are reported (%, maximum 100); higher is better. BenchGecko collects CSQA2 scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/csqa2 #### Cybench URL: https://benchgecko.ai/learn/glossary/cybench Built from BenchGecko data as of 2026-06-22 TL;DR: Cybench is a coding benchmark tracked on BenchGecko. Capture-the-flag cybersecurity challenges. Tests vulnerability analysis, reverse engineering, cryptography, and exploitation skills. Scores are reported (%, maximum 100); higher is better. BenchGecko collects Cybench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/cybench #### ANLI URL: https://benchgecko.ai/learn/glossary/anli Built from BenchGecko data as of 2026-10-02 TL;DR: ANLI is a knowledge benchmark tracked on BenchGecko. Adversarial Natural Language Inference. Tests ability to determine entailment, contradiction, or neutrality with adversarial examples. Scores are reported (%, maximum 100); higher is better. BenchGecko collects ANLI scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/anli #### The Agent Company URL: https://benchgecko.ai/learn/glossary/the-agent-company Built from BenchGecko data as of 2026-06-22 TL;DR: The Agent Company is an agentic benchmark tracked on BenchGecko. Simulated company environment testing agent ability to perform knowledge work tasks like research, analysis, and communication. Scores are reported (%, maximum 100); higher is better. BenchGecko collects The Agent Company scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/the-agent-company #### VideoMME URL: https://benchgecko.ai/learn/glossary/videomme Built from BenchGecko data as of 2026-04-09 TL;DR: VideoMME is a multimodal benchmark tracked on BenchGecko. Video understanding benchmark testing comprehension of video content, temporal reasoning, and scene analysis. Scores are reported (%, maximum 100); higher is better. BenchGecko collects VideoMME scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/videomme #### DeepResearch Bench URL: https://benchgecko.ai/learn/glossary/deepresearch-bench Built from BenchGecko data as of 2026-10-02 TL;DR: DeepResearch Bench is a knowledge benchmark tracked on BenchGecko. Tests ability to conduct in-depth research on complex topics, requiring synthesis of multiple sources. Scores are reported (%, maximum 100); higher is better. BenchGecko collects DeepResearch Bench scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/deepresearch-bench #### VPCT URL: https://benchgecko.ai/learn/glossary/vpct Built from BenchGecko data as of 2026-10-02 TL;DR: VPCT is a knowledge benchmark tracked on BenchGecko. Visual Pattern Completion Test. Tests abstract visual reasoning and pattern recognition ability. Scores are reported (%, maximum 100); higher is better. BenchGecko collects VPCT scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models. See live data: https://benchgecko.ai/benchmark/vpct ### Chips #### Trainium 3 URL: https://benchgecko.ai/learn/glossary/trainium-3 Text reviewed 2026-10-05 TL;DR: Trainium 3 is AWS's third-generation custom AI training chip, announced at re:Invent in December 2024. Trainium 3 targets training cost reduction for frontier models. The chip uses a NeuronCore architecture (custom AWS design, not CUDA-compatible). Software stack is AWS Neuron SDK + PyTorch XLA. See live data: https://benchgecko.ai/hardware #### Inferentia 3 URL: https://benchgecko.ai/learn/glossary/inferentia-3 Text reviewed 2026-10-05 TL;DR: AWS Inferentia is AWS's line of custom chips for AI inference · this entry covers the next generation after Inferentia2. Inferentia chips are used through AWS services such as EC2 Inf instances and Amazon Bedrock, with the AWS Neuron software stack. For training, AWS offers the separate Trainium line. See live data: https://benchgecko.ai/hardware #### AMD MI400 URL: https://benchgecko.ai/learn/glossary/mi400 Text reviewed 2026-10-05 TL;DR: AMD Instinct MI400 is the data-center GPU series AMD has said is planned for 2026, with HBM4 memory, following the MI300 and MI350 generations. MI400 is AMD's attempt to close the gap with NVIDIA on AI training. Roadmap context: MI300X (2023), MI325X (2024), MI400 (2026). Each generation targets 2× perf vs predecessor. HBM4 brings 1.5-2 TB/s bandwidth per stack, making memory-bound workloads (LLM inference, training gradient ops) faster. Packaging: CoWoS-L advanced packaging to increase die + HBM count. See live data: https://benchgecko.ai/hardware #### Microsoft Maia 200 URL: https://benchgecko.ai/learn/glossary/maia-200 Text reviewed 2026-10-05 TL;DR: Maia 200 is Microsoft's successor to Maia 100, its custom AI accelerator for Azure. BenchGecko publishes its figures once they are in the hardware dataset. Maia architecture: custom cores designed with OpenAI feedback, focused on large context (100K+) serving and reasoning workloads. Maia 200 specs leaked: HBM3e, 5nm, rack-scale via custom Cobalt-like interconnect. Software stack is an evolution of Microsoft's internal AI compiler (likely Triton-based). Primarily deployed inside Azure OpenAI · customers see it via lower prices on specific endpoints. See live data: https://benchgecko.ai/hardware #### H100 URL: https://benchgecko.ai/learn/glossary/h100 Built from BenchGecko data as of 2026-04-14 TL;DR: H100 is a NVIDIA GPU based on the Hopper SM90 architecture, first released in 2022. In the BenchGecko dataset it sits in the mainstream tier: current-generation non-flagship parts or last-generation flagships still widely deployed. It follows the A100 and is followed by the H200. Disclosed buyers include Meta and Microsoft. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/h100 #### H200 URL: https://benchgecko.ai/learn/glossary/h200 Built from BenchGecko data as of 2026-04-14 TL;DR: H200 is a NVIDIA GPU based on the Hopper SM90 architecture, first released in 2024. In the BenchGecko dataset it sits in the mainstream tier: current-generation non-flagship parts or last-generation flagships still widely deployed. It follows the H100. Disclosed buyers include Amazon. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/h200 #### B200 URL: https://benchgecko.ai/learn/glossary/b200 Built from BenchGecko data as of 2026-04-14 TL;DR: B200 is a NVIDIA GPU based on the Blackwell SM100 architecture, first released in 2024. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the H200 and is followed by the B300. Disclosed buyers include Meta, Microsoft and Oracle. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/b200 #### GB200 URL: https://benchgecko.ai/learn/glossary/gb200 Built from BenchGecko data as of 2026-04-14 TL;DR: GB200 is a NVIDIA CPU and GPU superchip based on the Grace CPU + 2x B200 architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It is followed by the GB300. Disclosed buyers include Oracle and CoreWeave. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/gb200 #### B300 URL: https://benchgecko.ai/learn/glossary/b300 Built from BenchGecko data as of 2026-04-14 TL;DR: B300 is a NVIDIA GPU based on the Blackwell Ultra SM100+ architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the B200. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/b300 #### GB300 URL: https://benchgecko.ai/learn/glossary/gb300 Built from BenchGecko data as of 2026-04-14 TL;DR: GB300 is a NVIDIA CPU and GPU superchip based on the Grace CPU + 2x B300 architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the GB200. Disclosed buyers include xAI. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/gb300 #### A100 URL: https://benchgecko.ai/learn/glossary/a100 Built from BenchGecko data as of 2026-04-14 TL;DR: A100 is a NVIDIA GPU based on the Ampere GA100 architecture, first released in 2020. In the BenchGecko dataset it sits in the legacy tier: end of life or being phased out. It is followed by the H100. Disclosed buyers include Meta. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/a100 #### MI300X URL: https://benchgecko.ai/learn/glossary/mi300x Built from BenchGecko data as of 2026-04-14 TL;DR: MI300X is an AMD GPU based on the CDNA 3 chiplet architecture, first released in 2023. In the BenchGecko dataset it sits in the mainstream tier: current-generation non-flagship parts or last-generation flagships still widely deployed. It is followed by the MI325X. Disclosed buyers include Microsoft and Meta. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/mi300x #### MI325X URL: https://benchgecko.ai/learn/glossary/mi325x Built from BenchGecko data as of 2026-04-14 TL;DR: MI325X is an AMD GPU based on the CDNA 3 chiplet architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the MI300X and is followed by the MI355X. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/mi325x #### MI355X URL: https://benchgecko.ai/learn/glossary/mi355x Built from BenchGecko data as of 2026-04-14 TL;DR: MI355X is an AMD GPU based on the CDNA 4 architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the MI325X. Disclosed buyers include Meta. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/mi355x #### TPU v5p URL: https://benchgecko.ai/learn/glossary/tpu-v5p Built from BenchGecko data as of 2026-04-14 TL;DR: TPU v5p is a Google TPU based on the Matrix MXU + VPU architecture, first released in 2024. In the BenchGecko dataset it sits in the mainstream tier: current-generation non-flagship parts or last-generation flagships still widely deployed. It is followed by the TPU v6e. Disclosed buyers include Google. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/tpu-v5p #### TPU v6e URL: https://benchgecko.ai/learn/glossary/tpu-v6e Built from BenchGecko data as of 2026-04-14 TL;DR: TPU v6e is a Google TPU based on the Matrix MXU + VPU + SparseCore architecture, first released in 2024. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the TPU v5p and is followed by the TPU v7. Disclosed buyers include Google. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/tpu-v6e #### TPU v7 URL: https://benchgecko.ai/learn/glossary/tpu-v7 Built from BenchGecko data as of 2026-04-14 TL;DR: TPU v7 is a Google TPU based on the Inference-first MXU architecture, first released in 2025. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. It follows the TPU v6e. Disclosed buyers include Google. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/tpu-v7 #### Trainium2 URL: https://benchgecko.ai/learn/glossary/trainium2 Built from BenchGecko data as of 2026-04-14 TL;DR: Trainium2 is an AWS custom AI chip (ASIC) based on the NeuronCore v3 architecture, first released in 2024. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. Disclosed buyers include Anthropic and Amazon. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/trainium2 #### MTIA v2 URL: https://benchgecko.ai/learn/glossary/mtia-v2 Built from BenchGecko data as of 2026-04-14 TL;DR: MTIA v2 is a Meta custom AI chip (ASIC) based on the PE grid + SRAM architecture, first released in 2024. In the BenchGecko dataset it sits in the specialized tier: wafer-scale, LPU and other custom designs. Disclosed buyers include Meta. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/mtia-v2 #### Gaudi 3 URL: https://benchgecko.ai/learn/glossary/gaudi-3 Built from BenchGecko data as of 2026-04-14 TL;DR: Gaudi 3 is an Intel custom AI chip (ASIC) based on the Heterogeneous MME + TPC architecture, first released in 2024. In the BenchGecko dataset it sits in the mainstream tier: current-generation non-flagship parts or last-generation flagships still widely deployed. Disclosed buyers include IBM. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/gaudi-3 #### Maia 100 URL: https://benchgecko.ai/learn/glossary/maia-100 Built from BenchGecko data as of 2026-04-14 TL;DR: Maia 100 is a Microsoft custom AI chip (ASIC) based on the Custom tensor core + MX format architecture, first released in 2024. In the BenchGecko dataset it sits in the specialized tier: wafer-scale, LPU and other custom designs. Disclosed buyers include Microsoft. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/maia-100 #### WSE-3 URL: https://benchgecko.ai/learn/glossary/wse-3 Built from BenchGecko data as of 2026-04-14 TL;DR: WSE-3 is a Cerebras wafer-scale AI processor based on the Wafer-scale SRAM mesh architecture, first released in 2024. In the BenchGecko dataset it sits in the specialized tier: wafer-scale, LPU and other custom designs. Disclosed buyers include G42. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/wse-3 #### LPU URL: https://benchgecko.ai/learn/glossary/groq-lpu Built from BenchGecko data as of 2026-04-14 TL;DR: LPU is a Groq custom AI chip (ASIC) based on the TSP deterministic dataflow architecture, first released in 2023. In the BenchGecko dataset it sits in the specialized tier: wafer-scale, LPU and other custom designs. Disclosed buyers include Aramco Digital. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/groq-lpu #### Ascend 910C URL: https://benchgecko.ai/learn/glossary/ascend-910c Built from BenchGecko data as of 2026-04-14 TL;DR: Ascend 910C is a Huawei custom AI chip (ASIC) based on the Da Vinci architecture, first released in 2024. In the BenchGecko dataset it sits in the frontier tier: current-generation flagship silicon shipping at scale to hyperscalers. Disclosed buyers include ByteDance and Baidu. Specs come from manufacturer datasheets; the spec card on this page shows the dataset values and their as-of date. See live data: https://benchgecko.ai/hardware/ascend-910c #### GPU URL: https://benchgecko.ai/learn/glossary/gpu TL;DR: A GPU is the accelerator that trains and serves most large AI models. A graphics processing unit executes thousands of parallel operations at once. Modern AI GPUs pair tensor cores with high bandwidth memory, making them the default hardware for training transformers and serving high-throughput inference. GPU supply, memory bandwidth, and networking are major constraints behind model pricing. ### Memory #### CoWoS-L URL: https://benchgecko.ai/learn/glossary/cowos-l Text reviewed 2026-10-05 TL;DR: CoWoS-L is TSMC's advanced packaging tech that uses a large local silicon interposer to pack more HBM and chiplets into a single GPU · foundation of Blackwell B300 and MI400. CoWoS-L uses Local Silicon Interconnect (LSI) bridges instead of a single monolithic interposer, lowering cost and enabling larger package sizes. Supports reticle-limit breaking designs (B300 is ~2× reticle size, physically impossible with CoWoS-S). Each wafer yields ~20-40 CoWoS packages depending on chip complexity. #### CoWoS-S URL: https://benchgecko.ai/learn/glossary/cowos-s TL;DR: CoWoS-S is the original Chip-on-Wafer-on-Substrate packaging · used on H100, A100 · being phased out for CoWoS-L in 2025-26. CoWoS-S uses a monolithic silicon interposer (reticle-size limited) vs CoWoS-L's Local Silicon Interconnect bridges. This limits package size and HBM count per chip. H100 with 80GB HBM3 is near the upper limit for CoWoS-S. Still used heavily in 2025-26 for H100 production; TSMC splits capacity between S and L based on customer demand. #### Hybrid Bonding URL: https://benchgecko.ai/learn/glossary/hybrid-bonding Text reviewed 2026-10-05 TL;DR: Hybrid bonding is a wafer-bonding technique that fuses two silicon wafers at the copper-to-copper level · critical for HBM4 and 3D stacked chips. Hybrid bonding requires extreme surface flatness (sub-nanometer) and purity · done in specialized cleanrooms. TSMC, Samsung, SK Hynix, and Intel all operate hybrid-bonding lines. HBM4 uses hybrid bonding to stack 16 DRAM dies (vs 12 in HBM3e) with wider interfaces. This pushes per-stack bandwidth from ~1.2 TB/s (HBM3e) to ~1.8-2 TB/s (HBM4). #### 3D NAND URL: https://benchgecko.ai/learn/glossary/3d-nand Text reviewed 2026-10-05 TL;DR: Vertically stacked NAND flash · used for AI training data storage · 300+ layers in 2026 flagship nodes. 3D NAND density doubles every 2-3 years via layer-count increases and smaller cell sizes. PCIe 5.0 + NVMe 2.0 SSDs based on 3D NAND can read at 12-14 GB/s · sufficient for training data ingestion. Enterprise flash arrays (Pure Storage, NetApp, VAST) use 3D NAND to build petabyte-scale low-latency storage that feeds GPU training jobs. 3D NAND is distinct from HBM (AI's fast memory) · NAND is bulk storage, HBM is on-package working memory. #### DDR6 URL: https://benchgecko.ai/learn/glossary/ddr6 Text reviewed 2026-10-05 TL;DR: DDR6 is the planned next generation of mainstream DRAM after DDR5, aimed at servers and the host CPUs that feed AI accelerators. DDR6 innovations: higher bus clock, increased burst length, on-die ECC default, and sub-channel architecture (two sub-channels per 64-bit channel) for better bandwidth utilization. Server DDR6 modules: RDIMM / LRDIMM / MCR-DIMM variants. Power-per-bit improved ~20% vs DDR5. CXL 3.0 integration allows DDR6 pools to attach to multiple CPUs. #### GDDR7 URL: https://benchgecko.ai/learn/glossary/gddr7 Built from BenchGecko data as of 2026-04-14 TL;DR: GDDR7 (Graphics DDR7) is a JEDEC GDDR memory generation from 2025, with 192 GB/s of bandwidth per device. Interface width is 256 bits at 1.1 V. Capacity per device: 2, 4 GB, up to 1-high stacks. Suppliers: Samsung (volume), SK hynix (ramping) and Micron (ramping). Main buyers: NVIDIA, AMD and Intel. #### HBM4 URL: https://benchgecko.ai/learn/glossary/hbm4 Built from BenchGecko data as of 2026-04-14 TL;DR: HBM4 (High Bandwidth Memory 4) is a JEDEC HBM memory generation from 2026, with 1,740 GB/s of bandwidth per stack. Interface width is 2,048 bits at 1.1 V. Capacity per stack: 36, 48 GB, up to 16-high stacks. Suppliers: SK hynix (sampling), Samsung (development) and Micron (development). Main buyers: NVIDIA, AMD and Google. See live data: https://benchgecko.ai/memory/hbm4 #### HBM3e URL: https://benchgecko.ai/learn/glossary/hbm3e Built from BenchGecko data as of 2026-04-14 TL;DR: HBM3e (High Bandwidth Memory 3e (Enhanced)) is a JEDEC HBM memory generation from 2024, with 1,180 GB/s of bandwidth per stack. Interface width is 1,024 bits at 1.1 V. Capacity per stack: 24, 36 GB, up to 12-high stacks. Suppliers: SK hynix (volume), Samsung (ramping) and Micron (volume). Main buyers: NVIDIA, AMD, Google, AWS, Microsoft and Meta. See live data: https://benchgecko.ai/memory/hbm3e #### HBM3 URL: https://benchgecko.ai/learn/glossary/hbm3 Built from BenchGecko data as of 2026-04-14 TL;DR: HBM3 (High Bandwidth Memory 3) is a JEDEC HBM memory generation from 2022, with 819 GB/s of bandwidth per stack. Interface width is 1,024 bits at 1.1 V. Capacity per stack: 16, 24 GB, up to 12-high stacks. Suppliers: SK hynix (volume), Samsung (volume) and Micron (volume). Main buyers: NVIDIA, AMD, Google and AWS. See live data: https://benchgecko.ai/memory/hbm3 #### HBM2e URL: https://benchgecko.ai/learn/glossary/hbm2e Built from BenchGecko data as of 2026-04-14 TL;DR: HBM2e (High Bandwidth Memory 2e (Enhanced)) is a JEDEC HBM memory generation from 2020, with 460 GB/s of bandwidth per stack. Interface width is 1,024 bits at 1.2 V. Capacity per stack: 8, 16 GB, up to 8-high stacks. Suppliers: Samsung (volume), SK hynix (volume) and Micron (eol). Main buyers: NVIDIA, Google and Intel. See live data: https://benchgecko.ai/memory/hbm2e #### DDR5 URL: https://benchgecko.ai/learn/glossary/ddr5 Built from BenchGecko data as of 2026-04-14 TL;DR: DDR5 (DDR5 SDRAM) is a JEDEC DDR memory generation from 2021, with 51 GB/s of bandwidth per device. Interface width is 64 bits at 1.1 V. Capacity per device: 8, 64 GB, up to 1-high stacks. Suppliers: Samsung (volume), SK hynix (volume), Micron (volume) and CXMT (ramping). Main buyers: Dell, HPE, Lenovo, Apple, Tesla and Intel. See live data: https://benchgecko.ai/memory/ddr5 #### LPDDR5X URL: https://benchgecko.ai/learn/glossary/lpddr5x Built from BenchGecko data as of 2026-04-14 TL;DR: LPDDR5X (Low Power DDR5X) is a JEDEC LPDDR memory generation from 2022, with 34 GB/s of bandwidth per device. Interface width is 64 bits at 0.5 V. Capacity per device: 8, 32 GB, up to 1-high stacks. Suppliers: Samsung (volume), SK hynix (volume) and Micron (volume). Main buyers: Apple, Samsung, Qualcomm and Google. See live data: https://benchgecko.ai/memory/lpddr5x #### GDDR6X URL: https://benchgecko.ai/learn/glossary/gddr6x Built from BenchGecko data as of 2026-04-14 TL;DR: GDDR6X (Graphics DDR6X) is a JEDEC GDDR memory generation from 2020, with 96 GB/s of bandwidth per device. Interface width is 256 bits at 1.35 V. Capacity per device: 1, 2 GB, up to 1-high stacks. Suppliers: Micron (volume) and Samsung (volume). Main buyers: NVIDIA. See live data: https://benchgecko.ai/memory/gddr6x ### Model families #### GPT (GPT) URL: https://benchgecko.ai/learn/glossary/gpt Text reviewed 2026-10-05 TL;DR: GPT is OpenAI's family of large language models · GPT stands for Generative Pre-trained Transformer · the models behind ChatGPT and the OpenAI API. GPT models are decoder-only transformers pre-trained on next-token prediction over large text and multimodal datasets, then post-trained with supervised fine-tuning and reinforcement learning from human and automated feedback. Recent generations combine general chat models and reasoning variants in one lineup, often with several sizes and price points per release. OpenAI also publishes some open-weight models; most GPT models are closed and available only through its API and products. See live data: https://benchgecko.ai/family/gpt #### Claude URL: https://benchgecko.ai/learn/glossary/claude Text reviewed 2026-10-05 TL;DR: Claude is Anthropic's family of large language models · first released in 2023 · offered in tiers from small and fast to large and most capable. Claude models are transformers post-trained with Anthropic's Constitutional AI method (published in 2022), in which the model critiques and revises its own outputs against a written set of principles, alongside reinforcement learning from feedback. Claude models support tool use and extended thinking, where the model reasons before answering and the thinking tokens are billed as output. Anthropic sells Claude through its own API and through cloud platforms such as Amazon Bedrock and Google Cloud Vertex AI. See live data: https://benchgecko.ai/family/claude #### Gemini URL: https://benchgecko.ai/learn/glossary/gemini Text reviewed 2026-10-05 TL;DR: Gemini is Google DeepMind's family of multimodal models · announced in December 2023 · the models behind the Gemini app and Google's AI products. Gemini models are trained on multimodal data from the start rather than adding vision to a text model afterwards. Google serves them through the Gemini API (Google AI Studio), Vertex AI on Google Cloud and its own products, on its TPU infrastructure. Thinking variants reason before answering, like other labs' reasoning models. See live data: https://benchgecko.ai/family/gemini #### Llama URL: https://benchgecko.ai/learn/glossary/llama Text reviewed 2026-10-05 TL;DR: Llama is Meta's family of open-weight large language models · first released in February 2023 · widely fine-tuned and self-hosted. Llama weights are released under the Llama Community License, which allows commercial use with conditions; it is not an OSI-approved open-source license, so "open-weight" is the accurate term. Because the weights are public, the same Llama model is hosted by many providers at different prices, and many community fine-tunes are built on it. See live data: https://benchgecko.ai/family/llama #### Qwen URL: https://benchgecko.ai/learn/glossary/qwen Text reviewed 2026-10-05 TL;DR: Qwen is Alibaba Cloud's family of large language models · first released in 2023 · a large set of open-weight models plus API-only flagships. The Qwen family includes chat, reasoning, coding, vision and audio variants. Qwen3 (2025) introduced hybrid models that can answer directly or think first. Because many sizes are open-weight, Qwen models are widely fine-tuned and hosted by third-party providers. See live data: https://benchgecko.ai/family/qwen #### DeepSeek URL: https://benchgecko.ai/learn/glossary/deepseek Text reviewed 2026-10-05 TL;DR: DeepSeek is a Chinese AI lab whose open-weight models, starting with DeepSeek-V3 (December 2024) and R1 (January 2025), showed frontier-class results at low prices. DeepSeek publishes technical reports with architecture and training details, which made its models a reference for the field: mixture-of-experts with many small experts, multi-head latent attention to shrink the KV cache, and FP8 training. R1 showed that reinforcement learning on tasks with checkable answers can produce long chain-of-thought reasoning. See live data: https://benchgecko.ai/family/deepseek #### Mistral URL: https://benchgecko.ai/learn/glossary/mistral Text reviewed 2026-10-05 TL;DR: Mistral is the model family of Mistral AI, a Paris-based lab founded in 2023 · a mix of open-weight models and commercial models served through its API. Mistral 7B used grouped-query attention and sliding-window attention to run efficiently. Mistral serves its models through its own platform (la Plateforme), its Le Chat assistant and cloud partners, and offers on-premise deployment for enterprises. See live data: https://benchgecko.ai/family/mistral ### Companies #### OpenAI URL: https://benchgecko.ai/learn/glossary/openai Built from BenchGecko data as of 2026-10-03 TL;DR: OpenAI is an AI model lab company based in San Francisco, CA, USA. Creator of ChatGPT and GPT series models, the world's most widely used AI platform. Key products: ChatGPT, GPT API, DALL-E, Sora and OpenAI Platform. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/openai #### Anthropic URL: https://benchgecko.ai/learn/glossary/anthropic Built from BenchGecko data as of 2026-10-03 TL;DR: Anthropic is an AI model lab company based in San Francisco, CA, USA. AI safety company building Claude, the leading enterprise AI assistant. Key products: Claude, Claude API, Claude Code and Claude for Enterprise. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/anthropic #### Google DeepMind URL: https://benchgecko.ai/learn/glossary/google-deepmind Built from BenchGecko data as of 2026-10-03 TL;DR: Google DeepMind is an AI model lab company based in London, UK. Alphabet's AI research lab, creator of Gemini models, AlphaFold, and other breakthrough AI systems. Key products: Gemini, Google AI Studio, Vertex AI, AlphaFold and Project Astra. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/google-deepmind #### Meta AI URL: https://benchgecko.ai/learn/glossary/meta-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Meta AI is an AI model lab company based in Menlo Park, CA, USA. Meta's AI research division. Key products: Llama models, Meta AI assistant, AI Studio and PyTorch. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/meta-ai #### xAI URL: https://benchgecko.ai/learn/glossary/xai Built from BenchGecko data as of 2026-10-03 TL;DR: xAI is an AI model lab company based in San Francisco, CA, USA. Elon Musk's AI company building Grok. Key products: Grok, Grok API, SuperGrok and Colossus (compute cluster). Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/xai #### Mistral AI URL: https://benchgecko.ai/learn/glossary/mistral-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Mistral AI is an AI model lab company based in Paris, France. Europe's leading AI company building open and commercial LLMs. Key products: Le Chat, La Plateforme API and Mistral models. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/mistral-ai #### Cohere URL: https://benchgecko.ai/learn/glossary/cohere Built from BenchGecko data as of 2026-10-03 TL;DR: Cohere is an AI model lab company based in Toronto, Canada. Enterprise-focused AI company providing NLP models for search, generation, and classification. Key products: Cohere API, Coral (chat), Command models and Embed models. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/cohere #### AI21 Labs URL: https://benchgecko.ai/learn/glossary/ai21-labs Built from BenchGecko data as of 2026-10-03 TL;DR: AI21 Labs is an AI model lab company based in Tel Aviv, Israel. Israeli AI company building enterprise-grade language models, including the Jamba architecture. Key products: AI21 Studio, Jamba models and Wordtune. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/ai21-labs #### Stability AI URL: https://benchgecko.ai/learn/glossary/stability-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Stability AI is an AI model lab company based in London, UK. Creator of Stable Diffusion and other open-source generative AI models for images, video, and audio. Key products: Stable Diffusion, Stable Video, Stable Audio and DreamStudio. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/stability-ai #### MiniMax URL: https://benchgecko.ai/learn/glossary/minimax Built from BenchGecko data as of 2026-10-03 TL;DR: MiniMax is an AI model lab company based in Shanghai, China. Chinese AI company building multimodal foundation models. Key products: Hailuo AI, Talkie and MiniMax API. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/minimax #### Zhipu AI URL: https://benchgecko.ai/learn/glossary/zhipu-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Zhipu AI is an AI model lab company based in Beijing, China. Chinese AI company behind GLM/ChatGLM models. Key products: Zhipu AI Platform, GLM models, CogView and CogVideo. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/zhipu-ai #### Baidu URL: https://benchgecko.ai/learn/glossary/baidu Built from BenchGecko data as of 2026-10-03 TL;DR: Baidu is an AI model lab company based in Beijing, China. Chinese tech giant and creator of ERNIE AI models. Key products: ERNIE Bot, Qianfan Platform, Baidu AI Cloud and Apollo Go (robotaxi). Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/baidu #### Alibaba (Qwen) URL: https://benchgecko.ai/learn/glossary/alibaba-qwen Built from BenchGecko data as of 2026-10-03 TL;DR: Alibaba (Qwen) is an AI model lab company based in Hangzhou, China. Chinese tech giant. Key products: Qwen models, Tongyi Qianwen, Alibaba Cloud AI and ModelScope. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/alibaba-qwen #### 01.AI URL: https://benchgecko.ai/learn/glossary/01-ai Built from BenchGecko data as of 2026-10-03 TL;DR: 01.AI is an AI model lab company based in Beijing, China. Kai-Fu Lee's AI startup building Yi models. Key products: Yi models and Yi Platform. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/01-ai #### Moonshot AI URL: https://benchgecko.ai/learn/glossary/moonshot-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Moonshot AI is an AI model lab company based in Beijing, China. Chinese AI company behind Kimi chatbot. Key products: Kimi and Kimi API. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/moonshot-ai #### StepFun URL: https://benchgecko.ai/learn/glossary/stepfun Built from BenchGecko data as of 2026-10-03 TL;DR: StepFun is an AI model lab company based in Shanghai, China. Chinese AI model maker founded by ex-Microsoft employees. Key products: StepChat, StepFun API and Step models. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/stepfun #### Xiaomi (MiMo) URL: https://benchgecko.ai/learn/glossary/xiaomi-mimo Built from BenchGecko data as of 2026-10-03 TL;DR: Xiaomi (MiMo) is an AI model lab company based in Beijing, China. Chinese tech giant's AI division building MiMo models for devices, agents, robots, and voice. Key products: MiMo models, Xiaomi HyperOS AI and Smart home AI. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/xiaomi-mimo #### Reka AI URL: https://benchgecko.ai/learn/glossary/reka-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Reka AI is an AI model lab company based in San Francisco, CA, USA. Multimodal AI startup building models that understand text, images, audio, and video natively. Key products: Reka models, Reka API and Reka Playground. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/reka-ai #### Inflection AI URL: https://benchgecko.ai/learn/glossary/inflection-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Inflection AI is an AI model lab company based in Palo Alto, CA, USA. AI company that built Pi chatbot. Key products: Inflection API and Pi (legacy, deprioritized). Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/inflection-ai #### Together AI URL: https://benchgecko.ai/learn/glossary/together-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Together AI is an AI compute infrastructure company based in San Francisco, CA, USA. AI cloud platform for training and running open-source models. Key products: Together Inference, Together Fine-tuning and Together GPU Clusters. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/together-ai #### Fireworks AI URL: https://benchgecko.ai/learn/glossary/fireworks-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Fireworks AI is an AI compute infrastructure company based in San Francisco, CA, USA. AI inference platform for deploying and serving models at scale. Key products: Fireworks Inference API, Fireworks Fine-tuning and FireFunction. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/fireworks-ai #### Groq URL: https://benchgecko.ai/learn/glossary/groq Built from BenchGecko data as of 2026-10-05 TL;DR: Groq is a semiconductor company based in Mountain View, CA, USA. AI inference chip company with LPU (Language Processing Unit). Key products: Groq LPU Inference Engine, GroqCloud and Groq 3 LPX (post-acquisition). Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/groq #### Replicate URL: https://benchgecko.ai/learn/glossary/replicate Built from BenchGecko data as of 2026-10-03 TL;DR: Replicate is an AI compute infrastructure company based in San Francisco, CA, USA. Platform for running ML models in the cloud via API. Key products: Replicate API, Replicate Predictions and Model hosting. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/replicate #### Modal URL: https://benchgecko.ai/learn/glossary/modal Built from BenchGecko data as of 2026-10-03 TL;DR: Modal is an AI compute infrastructure company based in New York, NY, USA. Serverless cloud for AI/ML workloads. Key products: Modal Cloud, Modal Functions and GPU scheduling. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/modal #### Anyscale URL: https://benchgecko.ai/learn/glossary/anyscale Built from BenchGecko data as of 2026-10-03 TL;DR: Anyscale is an AI compute infrastructure company based in San Francisco, CA, USA. Company behind Ray, the open-source framework for distributed AI. Key products: Anyscale Platform, Ray (open-source) and Anyscale Endpoints. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/anyscale #### Baseten URL: https://benchgecko.ai/learn/glossary/baseten Built from BenchGecko data as of 2026-10-03 TL;DR: Baseten is an AI compute infrastructure company based in San Francisco, CA, USA. AI inference infrastructure platform. Key products: Baseten Inference, Truss (open-source) and Model deployment. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/baseten #### Cerebras Systems URL: https://benchgecko.ai/learn/glossary/cerebras Built from BenchGecko data as of 2026-10-03 TL;DR: Cerebras Systems is a semiconductor company based in Sunnyvale, CA, USA. Wafer-scale AI chip company building the world's largest processors. Key products: Wafer-Scale Engine 3 (WSE-3), CS-3 system and Cerebras Inference. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/cerebras #### SambaNova Systems URL: https://benchgecko.ai/learn/glossary/sambanova Built from BenchGecko data as of 2026-10-03 TL;DR: SambaNova Systems is a semiconductor company based in Palo Alto, CA, USA. AI hardware company building custom chips for enterprise AI. Key products: SambaNova Suite, DataScale and SN40L chip. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/sambanova #### Lambda URL: https://benchgecko.ai/learn/glossary/lambda Built from BenchGecko data as of 2026-10-03 TL;DR: Lambda is an AI compute infrastructure company based in San Francisco, CA, USA. GPU cloud provider for AI training and inference. Key products: Lambda GPU Cloud, Lambda On-Demand, Lambda Reserved and Lambda Workstations. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/lambda #### Perplexity AI URL: https://benchgecko.ai/learn/glossary/perplexity Built from BenchGecko data as of 2026-10-03 TL;DR: Perplexity AI is an AI application company based in San Francisco, CA, USA. AI-powered answer engine replacing traditional search. Key products: Perplexity Search, Perplexity Pro, Perplexity API and Enterprise Perplexity. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/perplexity #### Character.AI URL: https://benchgecko.ai/learn/glossary/character-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Character.AI is an AI application company based in Menlo Park, CA, USA. AI chatbot platform for creating and conversing with AI characters. Key products: Character.AI Platform and Character creation tools. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/character-ai #### Jasper AI URL: https://benchgecko.ai/learn/glossary/jasper-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Jasper AI is an AI application company based in Austin, TX, USA. AI content creation platform for marketing teams. Key products: Jasper Platform, Jasper Brand Voice, Jasper Everywhere (browser ext) and Jasper API. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/jasper-ai #### Copy.ai URL: https://benchgecko.ai/learn/glossary/copy-ai Built from BenchGecko data as of 2026-10-05 TL;DR: Copy.ai is an AI application company based in San Francisco, CA, USA. AI-powered GTM platform for sales and marketing. Key products: Copy.ai Workflows, AI Sales OS and Content generation. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/copy-ai #### Runway URL: https://benchgecko.ai/learn/glossary/runway Built from BenchGecko data as of 2026-10-03 TL;DR: Runway is an AI application company based in New York, NY, USA. AI video generation and editing platform. Key products: Runway Gen-3, Runway Video Editor, Runway API and Runway Studios. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/runway #### Midjourney URL: https://benchgecko.ai/learn/glossary/midjourney Built from BenchGecko data as of 2026-10-03 TL;DR: Midjourney is an AI application company based in San Francisco, CA, USA. AI image generation platform. Key products: Midjourney image generation and Midjourney web app. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/midjourney #### ElevenLabs URL: https://benchgecko.ai/learn/glossary/elevenlabs Built from BenchGecko data as of 2026-10-03 TL;DR: ElevenLabs is an AI application company based in London, UK. AI voice synthesis and audio platform. Key products: ElevenLabs Voice AI, Text-to-Speech API, Voice cloning, AI Dubbing and Reader app. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/elevenlabs #### Synthesia URL: https://benchgecko.ai/learn/glossary/synthesia Built from BenchGecko data as of 2026-10-05 TL;DR: Synthesia is an AI application company based in London, UK. AI video platform for corporate training with AI-generated avatars. Key products: Synthesia Studio, AI Avatars, AI Video generation and Enterprise platform. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/synthesia #### Glean URL: https://benchgecko.ai/learn/glossary/glean Built from BenchGecko data as of 2026-10-03 TL;DR: Glean is an AI application company based in Palo Alto, CA, USA. AI-powered enterprise search and knowledge management platform. Key products: Glean Search, Glean Chat, Glean Agents and Glean Apps. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/glean #### Harvey AI URL: https://benchgecko.ai/learn/glossary/harvey-ai Built from BenchGecko data as of 2026-10-05 TL;DR: Harvey AI is an AI application company based in San Francisco, CA, USA. AI platform for law firms and legal professionals. Key products: Harvey for Law Firms, Harvey for Enterprise and AI Legal Agents. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/harvey-ai #### Anysphere (Cursor) URL: https://benchgecko.ai/learn/glossary/anysphere Built from BenchGecko data as of 2026-10-03 TL;DR: Anysphere (Cursor) is an AI developer tools company based in San Francisco, CA, USA. Creator of Cursor, the AI-first code editor. Key products: Cursor (AI code editor), Cursor Tab and Cursor Chat. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/anysphere #### Replit URL: https://benchgecko.ai/learn/glossary/replit Built from BenchGecko data as of 2026-10-03 TL;DR: Replit is an AI developer tools company based in San Francisco, CA, USA. AI-powered software development platform. Key products: Replit IDE, Replit Agent, Replit Deployments and Replit AI. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/replit #### Hugging Face URL: https://benchgecko.ai/learn/glossary/hugging-face Built from BenchGecko data as of 2026-10-03 TL;DR: Hugging Face is an AI developer tools company based in New York, NY, USA. The GitHub of machine learning. Key products: Hugging Face Hub, Transformers library, Inference API, Spaces and LeRobot. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/hugging-face #### LangChain URL: https://benchgecko.ai/learn/glossary/langchain Built from BenchGecko data as of 2026-10-03 TL;DR: LangChain is an AI developer tools company based in San Francisco, CA, USA. Framework for building LLM-powered applications. Key products: LangChain framework, LangSmith, LangGraph and LangServe. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/langchain #### Weights & Biases URL: https://benchgecko.ai/learn/glossary/weights-biases Built from BenchGecko data as of 2026-10-03 TL;DR: Weights & Biases is an AI developer tools company based in San Francisco, CA, USA. MLOps platform for experiment tracking and model management. Key products: W&B Experiment Tracking, W&B Sweeps, W&B Artifacts and W&B Launch. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/weights-biases #### Scale AI URL: https://benchgecko.ai/learn/glossary/scale-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Scale AI is an AI developer tools company based in San Francisco, CA, USA. AI data platform for training data, evaluation, and alignment. Key products: Scale Data Engine, Scale SEAL (evaluation), Scale Donovan (defense) and Scale Labs. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/scale-ai #### Labelbox URL: https://benchgecko.ai/learn/glossary/labelbox Built from BenchGecko data as of 2026-10-03 TL;DR: Labelbox is an AI developer tools company based in San Francisco, CA, USA. Data labeling and annotation platform for AI training. Key products: Labelbox Annotate, Labelbox Model and Labelbox Catalog. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/labelbox #### Snorkel AI URL: https://benchgecko.ai/learn/glossary/snorkel-ai Built from BenchGecko data as of 2026-10-03 TL;DR: Snorkel AI is an AI developer tools company based in Redwood City, CA, USA. Data-centric AI platform for programmatic labeling and data development. Key products: Snorkel Flow, Snorkel data labeling and AI data development. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/snorkel-ai #### Pinecone URL: https://benchgecko.ai/learn/glossary/pinecone Built from BenchGecko data as of 2026-10-03 TL;DR: Pinecone is an AI developer tools company based in New York, NY, USA. Managed vector database for AI applications. Key products: Pinecone vector database, Pinecone Serverless and Pinecone Assistants. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/pinecone #### Weaviate URL: https://benchgecko.ai/learn/glossary/weaviate Built from BenchGecko data as of 2026-10-03 TL;DR: Weaviate is an AI developer tools company based in Amsterdam, Netherlands. Open-source AI-native vector database. Key products: Weaviate vector database, Weaviate Cloud and Weaviate Hybrid Search. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/weaviate #### NVIDIA URL: https://benchgecko.ai/learn/glossary/nvidia Built from BenchGecko data as of 2026-10-03 TL;DR: NVIDIA is a semiconductor company based in Santa Clara, CA, USA. NVIDIA's products include Blackwell GPUs, CUDA, TensorRT, NIM, DGX Cloud and Groq 3 LPX. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/nvidia #### Microsoft URL: https://benchgecko.ai/learn/glossary/microsoft Built from BenchGecko data as of 2026-10-03 TL;DR: Microsoft is a big tech company based in Redmond, WA, USA. Major AI investor (OpenAI, Anthropic). Key products: Azure AI, Microsoft Copilot, Azure OpenAI Service, GitHub Copilot and Phi models. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/microsoft #### Amazon (AWS) URL: https://benchgecko.ai/learn/glossary/amazon-aws Built from BenchGecko data as of 2026-10-03 TL;DR: Amazon (AWS) is a big tech company based in Seattle, WA, USA. Leading cloud provider. Key products: AWS Bedrock, Amazon Nova, Amazon Q, Trainium chips and SageMaker. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/amazon-aws #### CoreWeave URL: https://benchgecko.ai/learn/glossary/coreweave Built from BenchGecko data as of 2026-10-03 TL;DR: CoreWeave is an AI compute infrastructure company based in Livingston, NJ, USA. GPU cloud infrastructure provider for AI workloads. Key products: CoreWeave Cloud, GPU-as-a-Service and Kubernetes GPU clusters. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/coreweave #### Databricks URL: https://benchgecko.ai/learn/glossary/databricks Built from BenchGecko data as of 2026-10-03 TL;DR: Databricks is an AI compute infrastructure company based in San Francisco, CA, USA. Data and AI platform company. Key products: Databricks Lakehouse, Mosaic AI, Lakebase, Unity Catalog and Delta Lake. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/databricks #### Alphabet (Google) URL: https://benchgecko.ai/learn/glossary/google Built from BenchGecko data as of 2026-10-03 TL;DR: Alphabet (Google) is a big tech company based in Mountain View, CA, USA. Parent of Google DeepMind. Key products: Google Cloud AI, Gemini, Vertex AI, Google AI Studio and Gemini API. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/google #### Meta Platforms URL: https://benchgecko.ai/learn/glossary/meta Built from BenchGecko data as of 2026-10-03 TL;DR: Meta Platforms is a big tech company based in Menlo Park, CA, USA. Social media giant. Key products: Llama models, Meta AI, AI Studio, PyTorch and WhatsApp AI Agents. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/meta #### IREN URL: https://benchgecko.ai/learn/glossary/iren Built from BenchGecko data as of 2026-10-03 TL;DR: IREN is an energy company based in Sydney, Australia. AI datacenter and Bitcoin mining operator. Key products: AI Cloud GPU Hosting, Bitcoin Mining and Data Center Operations. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/iren #### Nebius URL: https://benchgecko.ai/learn/glossary/nebius Built from BenchGecko data as of 2026-10-03 TL;DR: Nebius is an AI compute infrastructure company based in Amsterdam, Netherlands. AI infrastructure company spun off from Yandex. Key products: Nebius AI Cloud, GPU Clusters and ML Platform. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/nebius #### Applied Digital URL: https://benchgecko.ai/learn/glossary/applied-digital Built from BenchGecko data as of 2026-10-03 TL;DR: Applied Digital is an energy company based in Dallas, TX, USA. AI datacenter operator building high-performance computing facilities in Texas and North Dakota. Key products: HPC Datacenter Hosting, AI Cloud Services and Liquid Cooling Infrastructure. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/applied-digital #### Crusoe Energy URL: https://benchgecko.ai/learn/glossary/crusoe-energy Built from BenchGecko data as of 2026-10-03 TL;DR: Crusoe Energy is an energy company based in Denver, CO, USA. AI compute powered by stranded energy. Key products: Crusoe Cloud, Stranded Gas GPU Compute and Wind-Powered AI Datacenter. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/crusoe-energy #### TSMC URL: https://benchgecko.ai/learn/glossary/tsmc Built from BenchGecko data as of 2026-10-03 TL;DR: TSMC is a chip foundry company based in Hsinchu, Taiwan. The world's largest semiconductor foundry. Key products: N3 Process, N2 Process, CoWoS Packaging, InFO Packaging and SoIC. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/tsmc #### Samsung Electronics URL: https://benchgecko.ai/learn/glossary/samsung-electronics Built from BenchGecko data as of 2026-10-03 TL;DR: Samsung Electronics is a memory chip company based in Suwon, South Korea. Global leader in memory semiconductors (DRAM, NAND, HBM) and second-largest foundry. Key products: HBM3E, LPDDR5X, GDDR7, 3nm Gate-All-Around and V-NAND. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/samsung-electronics #### SK Hynix URL: https://benchgecko.ai/learn/glossary/sk-hynix Built from BenchGecko data as of 2026-10-03 TL;DR: SK Hynix is a memory chip company based in Icheon, South Korea. Second-largest memory chipmaker and dominant HBM supplier. Key products: HBM3E 12-Hi, HBM4, DDR5, LPDDR5X and NAND. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/sk-hynix #### Micron Technology URL: https://benchgecko.ai/learn/glossary/micron Built from BenchGecko data as of 2026-10-03 TL;DR: Micron Technology is a memory chip company based in Boise, ID, USA. Third-largest memory chipmaker. Key products: HBM3E, DDR5, LPDDR5X, GDDR7 and QLC NAND. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/micron #### ASML URL: https://benchgecko.ai/learn/glossary/asml Built from BenchGecko data as of 2026-10-03 TL;DR: ASML is a chip foundry company based in Veldhoven, Netherlands. Sole manufacturer of EUV lithography machines used to fabricate all advanced AI chips. Key products: EUV Lithography, High-NA EUV, DUV Lithography, Metrology and YieldStar. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/asml #### Broadcom URL: https://benchgecko.ai/learn/glossary/broadcom Built from BenchGecko data as of 2026-10-03 TL;DR: Broadcom is a semiconductor company based in Palo Alto, CA, USA. Designs custom AI accelerators (TPUs for Google, Trainium for Amazon) and dominates AI networking (Tomahawk, Jericho switches). Key products: Custom AI ASICs, Tomahawk 5, Jericho3-AI, VMware Cloud Foundation and Ethernet Fabric. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/broadcom #### AMD URL: https://benchgecko.ai/learn/glossary/amd Built from BenchGecko data as of 2026-10-03 TL;DR: AMD is a semiconductor company based in Santa Clara, CA, USA. Second-largest GPU maker for AI training and inference. Key products: MI300X, MI325X, MI350, EPYC CPUs and Ryzen AI. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/amd #### Intel URL: https://benchgecko.ai/learn/glossary/intel Built from BenchGecko data as of 2026-10-03 TL;DR: Intel is a semiconductor company based in Santa Clara, CA, USA. Intel's products include Gaudi 3, Intel 18A, Xeon CPUs, Arc GPUs and Foundry Services. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/intel #### Qualcomm URL: https://benchgecko.ai/learn/glossary/qualcomm Built from BenchGecko data as of 2026-10-03 TL;DR: Qualcomm is a semiconductor company based in San Diego, CA, USA. Dominates mobile AI chips (Snapdragon). Key products: Snapdragon X Elite, Snapdragon 8 Gen 4, Cloud AI 100 and Automotive Ride Platform. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/qualcomm #### Marvell Technology URL: https://benchgecko.ai/learn/glossary/marvell Built from BenchGecko data as of 2026-10-03 TL;DR: Marvell Technology is a semiconductor company based in Wilmington, DE, USA. Designs custom AI silicon and electro-optics for hyperscale data centers. Key products: Custom AI Accelerators, PAM4 DSPs, PCIe 6.0 Switches and Coherent Optics. Valuation, revenue and funding are tracked with dates and sources on the company page; the snapshot on this page shows the current figures. See live data: https://benchgecko.ai/economy/company/marvell ### Foundry #### Samsung Foundry URL: https://benchgecko.ai/learn/glossary/samsung-foundry Built from BenchGecko data as of 2026-04-14 TL;DR: Samsung Foundry is a semiconductor foundry headquartered in South Korea that manufactures chips for other companies. Process nodes in the dataset: 4LPP (volume, 2023), 3GAA (ramping, 2024), 2GAA (development, 2026) and SF5 (volume, 2021). Clients include Qualcomm, Google and Tenstorrent. See live data: https://benchgecko.ai/foundries/samsung-foundry #### Intel Foundry URL: https://benchgecko.ai/learn/glossary/intel-foundry Built from BenchGecko data as of 2026-04-14 TL;DR: Intel Foundry is a semiconductor foundry headquartered in United States that manufactures chips for other companies. Process nodes in the dataset: Intel 4 (volume, 2023), Intel 3 (volume, 2024), Intel 18A (risk, 2025) and Intel 20A (limited, 2024). Clients include Microsoft and Intel (internal). See live data: https://benchgecko.ai/foundries/intel-foundry #### SMIC URL: https://benchgecko.ai/learn/glossary/smic Built from BenchGecko data as of 2026-04-14 TL;DR: SMIC is a semiconductor foundry headquartered in China that manufactures chips for other companies. Process nodes in the dataset: N7 (DUV) (volume, 2023), N5 (DUV) (development, 2025), N14 (volume, 2019) and N28 (mature, 2015). Clients include Huawei and Biren. See live data: https://benchgecko.ai/foundries/smic #### GlobalFoundries URL: https://benchgecko.ai/learn/glossary/globalfoundries Built from BenchGecko data as of 2026-04-14 TL;DR: GlobalFoundries is a semiconductor foundry headquartered in United States that manufactures chips for other companies. Process nodes in the dataset: 14LPP (volume, 2016), 12LP (volume, 2018) and 22FDX (volume, 2019). Clients include Groq, AMD (legacy) and US DoD. See live data: https://benchgecko.ai/foundries/globalfoundries ### Systems #### DGX GB200 NVL72 URL: https://benchgecko.ai/learn/glossary/nvidia-dgx-gb200-nvl72 Built from BenchGecko data as of 2026-04-14 TL;DR: DGX GB200 NVL72 is a NVIDIA AI system with 72 NVIDIA B200 and 36 NVIDIA Grace. The dataset lists 720 PFLOPS of FP8 compute and 360 PFLOPS at BF16. It has 13.8 TB of HBM3e. Interconnect: NVLink 5.0. Power: 120 kW, liquid cooled. See live data: https://benchgecko.ai/systems/nvidia-dgx-gb200-nvl72 #### DGX B200 URL: https://benchgecko.ai/learn/glossary/nvidia-dgx-b200 Built from BenchGecko data as of 2026-04-14 TL;DR: DGX B200 is a NVIDIA AI system with 8 NVIDIA B200 and 2 Dual Intel Xeon. The dataset lists 80 PFLOPS of FP8 compute and 40 PFLOPS at BF16. It has 1.5 TB of HBM3e. Interconnect: NVLink 4.0. Power: 14 kW, air cooled. See live data: https://benchgecko.ai/systems/nvidia-dgx-b200 #### HGX H100 URL: https://benchgecko.ai/learn/glossary/nvidia-hgx-h100 Built from BenchGecko data as of 2026-04-14 TL;DR: HGX H100 is a NVIDIA AI system with 8 NVIDIA H100 SXM and 2 Dual Intel Xeon. The dataset lists 32 PFLOPS of FP8 compute and 16 PFLOPS at BF16. It has 0.6 TB of HBM3. Interconnect: NVLink 4.0. Power: 10 kW, air cooled. See live data: https://benchgecko.ai/systems/nvidia-hgx-h100 #### TPU v6e Pod URL: https://benchgecko.ai/learn/glossary/google-tpu-v6e-pod Built from BenchGecko data as of 2026-04-14 TL;DR: TPU v6e Pod is a Google AI system with 256 Google TPU v6e and 64 Custom host. The dataset lists 230 PFLOPS of FP8 compute and 115 PFLOPS at BF16. It has 8 TB of HBM3e. Interconnect: ICI 4.0. Power: 60 kW, liquid cooled. See live data: https://benchgecko.ai/systems/google-tpu-v6e-pod #### MI325X Platform URL: https://benchgecko.ai/learn/glossary/amd-instinct-mi325x-platform Built from BenchGecko data as of 2026-04-14 TL;DR: MI325X Platform is an AMD AI system with 8 AMD MI325X and 2 Dual AMD EPYC. The dataset lists 48 PFLOPS of FP8 compute and 24 PFLOPS at BF16. It has 2 TB of HBM3e. Interconnect: Infinity Fabric. Power: 10 kW, air cooled. See live data: https://benchgecko.ai/systems/amd-instinct-mi325x-platform #### CS-3 URL: https://benchgecko.ai/learn/glossary/cerebras-cs3 Built from BenchGecko data as of 2026-04-14 TL;DR: CS-3 is a Cerebras AI system with 1 Cerebras WSE-3 and 1 Host server. The dataset lists 125 PFLOPS at BF16. Interconnect: SwarmX. Power: 23 kW, liquid cooled. See live data: https://benchgecko.ai/systems/cerebras-cs3 #### DGX GB300 NVL72 URL: https://benchgecko.ai/learn/glossary/nvidia-dgx-gb300-nvl72 Built from BenchGecko data as of 2026-04-14 TL;DR: DGX GB300 NVL72 is a NVIDIA AI system with 72 NVIDIA B300 and 36 NVIDIA Grace Ultra. The dataset lists 1,440 PFLOPS of FP8 compute and 720 PFLOPS at BF16. It has 20.7 TB of HBM3e. Interconnect: NVLink 5.0+. Power: 140 kW, liquid cooled. See live data: https://benchgecko.ai/systems/nvidia-dgx-gb300-nvl72 #### TPU v5p Pod URL: https://benchgecko.ai/learn/glossary/google-tpu-v5p-pod Built from BenchGecko data as of 2026-04-14 TL;DR: TPU v5p Pod is a Google AI system with 8,960 Google TPU v5p and 2,240 Custom host. The dataset lists 8,100 PFLOPS of FP8 compute and 4,050 PFLOPS at BF16. It has 860 TB of HBM3. Interconnect: ICI 3.0. Power: 2,500 kW, liquid cooled. See live data: https://benchgecko.ai/systems/google-tpu-v5p-pod #### Trn2 UltraServer URL: https://benchgecko.ai/learn/glossary/aws-trainium2-ultraserver Built from BenchGecko data as of 2026-04-14 TL;DR: Trn2 UltraServer is an AWS AI system with 16 AWS Trainium2 and 2 AWS Graviton4. The dataset lists 48 PFLOPS of FP8 compute and 24 PFLOPS at BF16. It has 1.5 TB of HBM3. Interconnect: NeuronLink. Power: 12 kW, air cooled. See live data: https://benchgecko.ai/systems/aws-trainium2-ultraserver #### Maia 100 Rack URL: https://benchgecko.ai/learn/glossary/microsoft-maia-100-rack Built from BenchGecko data as of 2026-04-14 TL;DR: Maia 100 Rack is a Microsoft AI system with 32 Microsoft Maia 100 and 8 Microsoft Cobalt 100. The dataset lists 96 PFLOPS of FP8 compute and 48 PFLOPS at BF16. It has 6.1 TB of HBM3e. Interconnect: Custom Ethernet fabric. Power: 40 kW, liquid cooled. See live data: https://benchgecko.ai/systems/microsoft-maia-100-rack ### Concepts #### Mixture of Experts (MoE) URL: https://benchgecko.ai/learn/glossary/moe Text reviewed 2026-10-05 TL;DR: A model architecture where only a subset of experts activate per token, slashing inference cost while preserving quality. In a dense transformer, every parameter fires for every token. In a Mixture-of-Experts transformer, the feedforward layer is replaced by a bank of experts and a router. For each token the router picks the top-k experts (usually 2 out of 64 or more) and only those experts compute. The trade-offs: load balancing loss to keep experts busy, routing instability during training, and higher total VRAM at serving time because all experts must be resident even if only a few are used per token. See the related terms and live BenchGecko data for current examples. #### Reasoning Model URL: https://benchgecko.ai/learn/glossary/reasoning-model Text reviewed 2026-10-05 TL;DR: A model that generates internal reasoning tokens before producing the answer, trading inference cost for accuracy on math and logic. Reasoning models are trained with reinforcement learning on chain-of-thought traces that maximize a correctness reward. At inference time they allocate a variable "thinking budget" of tokens that are generated but not shown to the user. OpenAI o1 spends between 1K and 100K reasoning tokens per problem depending on difficulty. Claude Extended Thinking exposes the reasoning transparently and charges for it. DeepSeek R1 open-sourced both the weights and the RL recipe, making reasoning models accessible to anyone. The trade-off is direct: more thinking tokens equals higher cost and latency but better accuracy on benchmarks that require multi-step inference. For pure recall or chat, reasoning models are overkill. #### RAG (RAG) URL: https://benchgecko.ai/learn/glossary/rag TL;DR: A technique that retrieves relevant documents before answering, so the model grounds its output in real data instead of fabricating. A RAG system has three components: an embedding model (e.g., text-embedding-3-large, Voyage, Cohere embed), a vector database (Pinecone, Weaviate, pgvector, Chroma), and a generation model (any LLM). Query → embedding → similarity search → top-k documents → prompt template with retrieved context → generation. Quality hinges on: chunk size, chunk overlap, retrieval ranking, reranking (Cohere Rerank, Voyage Rerank), and prompt-template design. Advanced patterns include: hybrid search (sparse + dense), HyDE (hypothetical document embeddings), query decomposition, and graph RAG. Cost dominates on embedding (re-embedding every new document) and inference (larger contexts from retrieved docs). #### MCP (MCP) URL: https://benchgecko.ai/learn/glossary/mcp Text reviewed 2026-10-05 TL;DR: An open protocol from Anthropic that standardizes how AI models talk to external tools, data sources, and services. MCP uses JSON-RPC over stdio, SSE, or WebSocket. Servers expose three primitives: resources (readable state), tools (callable functions), and prompts (predefined templates). Clients connect to servers at startup, list capabilities, and surface them to the model. The model decides when to invoke tools; the client executes and returns results. MCP solves the N×M integration problem: N AI applications × M tools collapse to N + M MCP implementations. The protocol is spec-first and open; Anthropic published the SDK in multiple languages. Cursor, Cline, Claude Desktop, Zed, and most AI coding agents now ship MCP support. #### Subagent URL: https://benchgecko.ai/learn/glossary/subagent TL;DR: A subagent is a child agent spawned by a parent to handle a scoped sub-task · common pattern in Claude Code, Roo Code, and multi-agent frameworks. Subagent systems solve two problems: context window limits (each subagent has fresh context) and task specialization (each gets tailored prompts and tool subsets). Coordination patterns vary: Claude Code uses a parent that issues Task tool calls, each spawning a subagent that returns a summary. Roo Code runs multi-tab orchestration with persistent sub-workspaces. AutoGen and CrewAI formalize the pattern with role-named agents and turn-taking. #### Sliding Window Attention URL: https://benchgecko.ai/learn/glossary/sliding-window-attention Text reviewed 2026-10-05 TL;DR: See the related terms and live BenchGecko data for current examples. Sliding window works because most useful context is local. Information that needs to travel far uses the "receptive field" expansion through stacked layers · a 32-layer model with 4K window has effective context of 128K through depth. Compute cost: sliding window attention is ~K/N cheaper than full attention when N >> K. See the related terms and live BenchGecko data for current examples. #### ALiBi (ALiBi) URL: https://benchgecko.ai/learn/glossary/alibi TL;DR: ALiBi (Attention with Linear Biases) replaces positional embeddings with a linear bias added to attention scores · used in BLOOM, MPT, some open-source models. ALiBi bias: attention_score(i,j) = qk(i,j) + m × |i - j|, where m is a head-specific negative constant. Heads get different m values, so some heads look far (small m) and some close (large m). Key property: ALiBi models extrapolate · trained on 2K context, they work decently on 8K+ without fine-tuning. This made ALiBi popular for researchers who wanted to test long-context without retraining. #### Mamba URL: https://benchgecko.ai/learn/glossary/mamba Text reviewed 2026-10-05 TL;DR: Mamba is a state-space model (SSM) architecture that scales linearly with sequence length · 5-10× faster than transformers at long context · powers Jamba, Zamba, Mamba-2. Mamba's core innovation is making state-space models content-aware · the state update depends on the input token, so the model can "selectively forget" or retain. This solves earlier SSM limitations (previous SSMs like S4 were linear, lost expressiveness). Mamba models perform competitively with transformers at small scales but quality gap at 70B+ remains unclear. Hybrid models like Jamba (AI21) mix Mamba + transformer layers to get best of both. #### Hyena URL: https://benchgecko.ai/learn/glossary/hyena TL;DR: Hyena is an implicit-convolution operator that replaces attention · scales O(N log N) · studied by Together AI and used in Evo (biology foundation model). Hyena operator: learn a function f(t) that generates convolution kernels on-the-fly based on relative position. This implicit convolution is computed in frequency domain via FFT, giving O(N log N) complexity. Together AI's Evo (131K context, biology) showed Hyena works on long-sequence biology tasks. For text, Hyena has seen less adoption than Mamba · the architecture is harder to tune and lacks Mamba's selective-state flexibility. #### KV Cache Compression URL: https://benchgecko.ai/learn/glossary/kv-cache-compression Text reviewed 2026-10-05 TL;DR: KV cache compression reduces the memory cost of long-context LLM inference · via quantization, eviction, or sparsity · cuts VRAM 2-8× for long contexts. Techniques: (1) KV quantization (INT4/INT8 stored KV tensors) cuts memory 2-4× with minor quality loss. (2) Eviction policies (H2O, StreamingLLM) drop low-attention tokens. (3) Sparsity (e.g., MQA, GQA at inference) reduce K+V channels. (4) Prefix caching (Anthropic, OpenAI) reuses KV across sessions · not compression per se but related. vLLM, SGLang, TensorRT-LLM all ship KV cache management · critical for production long-context serving. #### Fine-tuning URL: https://benchgecko.ai/learn/glossary/fine-tuning Text reviewed 2026-10-05 TL;DR: Training a base model on a smaller, specific dataset to teach it a new skill or voice. Fine-tuning recipes: Supervised Fine-Tuning (SFT) on labeled demonstrations · the most common. Direct Preference Optimization (DPO) on preference pairs · preferred over RLHF for alignment because it skips reward modeling. RLHF still used for safety-critical alignment. LoRA (Low-Rank Adaptation) trains a small adapter matrix while freezing base weights · 10-100× cheaper than full fine-tune, recoverable to base. QLoRA adds 4-bit quantization to LoRA for consumer-GPU training. Fine-tuning is reversible by discarding the adapter. See the related terms and live BenchGecko data for current examples. #### Chain of Thought (CoT) URL: https://benchgecko.ai/learn/glossary/chain-of-thought TL;DR: Prompting the model to show its reasoning steps before the final answer · dramatically improves math, logic, and multi-step tasks. CoT exploits the transformer's auto-regressive generation: each token conditions on all previous tokens. By generating intermediate reasoning, the model has more compute per final-answer token and can catch errors in earlier steps. Zero-shot CoT ("Let's think step by step") was the original trigger. Few-shot CoT provides worked examples in the prompt. Models trained on CoT data (Chinchilla and later) produce CoT naturally. Self-consistency improves CoT by sampling multiple reasoning paths and voting on the final answer. Tree-of-Thought extends CoT to branching exploration. For tasks with verifiable answers, sample-and-verify schemes dramatically outperform single-shot CoT. #### Quantization URL: https://benchgecko.ai/learn/glossary/quantization TL;DR: Compressing model weights to lower-precision numbers (INT8, INT4, FP8) to cut memory use and speed up inference. Post-Training Quantization (PTQ): apply after training, fast. Methods: GPTQ (second-order error minimization), AWQ (activation-aware), SmoothQuant (migrates outliers). Quantization-Aware Training (QAT): train with quantized forward, backward in float. Slower but better quality. Weight-only vs weight+activation: weight-only is more forgiving; activation quantization requires calibration. Typical quality loss for 4-bit: 1-3 points on benchmarks. For 2-bit: 5-15 points. Lookup-based methods (GGUF k-quants) mix precision across layers. Frontier serving uses FP8 for both weight and activation to leverage Tensor Core acceleration. #### Distillation URL: https://benchgecko.ai/learn/glossary/distillation Text reviewed 2026-10-05 TL;DR: Training a smaller student model to reproduce a larger teacher model's outputs · cheaper to serve with similar quality. Teacher-student distillation: the student minimizes KL divergence between its output distribution and the teacher's on a training set. Temperature softens the teacher distribution · higher temperatures surface information about non-top tokens. Distillation targets: logits (most common), hidden states (feature matching), attention maps (relation-matching), or generated sequences (sequence-level). For large-scale distillation of frontier models, sequence-level works best · train the student on teacher-generated outputs as if they were ground truth. Cost: producing a high-quality distillation corpus is expensive (you pay to run the teacher at scale), but once done, student training is cheap. #### Context Window URL: https://benchgecko.ai/learn/glossary/context-window Text reviewed 2026-10-05 TL;DR: The max number of tokens · input + output · a model can handle in a single request. A context window includes the system prompt, user prompt, any retrieved documents (RAG), tool call results, prior conversation turns, and the generated output. Different models report context in different ways: effective context (where quality is preserved) is often smaller than maximum context. "Needle in a haystack" tests measure retrieval quality at depth · most models degrade past 50-70% depth. Compute cost of attention is quadratic in sequence length O(n²), but modern serving uses optimized attention (FlashAttention, ring attention) to push this lower. Pricing is typically linear per-token on input and a multiplier on output. #### Transformer URL: https://benchgecko.ai/learn/glossary/transformer TL;DR: The 2017 "Attention Is All You Need" architecture · parallelizable, scalable, and the foundation of every modern LLM. A transformer block has: multi-head self-attention, layer normalization, a feedforward MLP, and residual connections. "Attention" computes query-key-value projections of each token, takes dot products of queries with all keys, softmaxes them into weights, and outputs a weighted sum of values. Multi-head attention runs several of these in parallel at different subspaces, capturing different relational patterns. Positional embeddings (absolute, RoPE, ALiBi) encode token position since attention itself is order-agnostic. Decoder-only transformers (GPT family, Llama, Claude) are the dominant architecture for generative LLMs. MoE transformers replace the feedforward with a bank of experts plus a router (see MoE entry). #### AI Agent URL: https://benchgecko.ai/learn/glossary/agent Text reviewed 2026-10-05 TL;DR: An AI that plans multi-step workflows, uses tools, and maintains state · not a single-turn chat. Agent architectures vary: ReAct (reason + act interleaved), function-calling loops, tree-search planners, and hybrid designs. The agent runs in a loop: reason about next step → call tool → observe result → reason again → ... until done. State management includes conversation history, file system state, and task decomposition. Safety-critical: agents can make irreversible changes (send emails, modify files, spend money). Most production agents require explicit user approval for high-risk actions. See the related terms and live BenchGecko data for current examples. #### Tokens URL: https://benchgecko.ai/learn/glossary/tokens Text reviewed 2026-10-05 TL;DR: The fundamental unit that LLMs read and generate · 1 token ≈ 0.75 English words or 4 characters. Tokenization is language-specific. English averages 1.3 tokens per word. Code averages 2.5. Chinese and other non-Latin scripts average 2-4× more tokens per character than English, which is why non-English API calls cost more. Tokenizer algorithms vary: GPT uses BPE (Byte Pair Encoding), Llama uses SentencePiece, Claude uses a Claude-specific tokenizer. Larger vocab means fewer tokens per input but larger embedding tables. See the related terms and live BenchGecko data for current examples. #### Inference URL: https://benchgecko.ai/learn/glossary/inference TL;DR: The process of running a trained model to generate predictions · every API call is inference. Inference has two phases: prefill (process the input prompt in parallel) and decode (generate output tokens one at a time). Prefill is compute-bound; decode is memory-bandwidth-bound due to KV cache reads. This is why long outputs are expensive: each token requires a full pass reading the cache. Optimizations include: PagedAttention (efficient KV cache), speculative decoding (draft model predicts tokens, verified by main model), continuous batching (pack requests dynamically), and quantization (INT8/FP8 weights + activations). #### Latency URL: https://benchgecko.ai/learn/glossary/latency Text reviewed 2026-10-05 TL;DR: Time from request sent to first response token · the number that determines UX feel. TTFT depends on: queue wait (how many requests are ahead), prefill time (proportional to input length), model size (bigger = slower), and geographic distance (speed of light from your server to theirs). Reasoning models add 30-120s thinking latency before the first visible token. Groq and Cerebras hit sub-100ms TTFT by using exotic hardware (LPU, wafer-scale). Most frontier APIs: 400-2000ms TTFT at p50. #### Hallucination URL: https://benchgecko.ai/learn/glossary/hallucination Text reviewed 2026-10-05 TL;DR: When a model generates confident-sounding output that is factually wrong, fabricated, or unsupported by reality. Hallucination root cause: LLMs predict next-token probability, not truth. When the model lacks information, it predicts plausible continuations based on pattern-matching, producing output that looks right but isn't. Older models: 10-20%+. Mitigations rank: RAG > reasoning + verification > RLHF > prompt engineering. Hallucination is hardest in: obscure facts, numeric details, citations, legal/medical advice, code that uses deprecated APIs. See the related terms and live BenchGecko data for current examples. #### Scaling Laws URL: https://benchgecko.ai/learn/glossary/scaling-laws Text reviewed 2026-10-05 TL;DR: Empirical laws showing model quality improves predictably with more compute, parameters, and data · the foundation of modern AI scaling. Kaplan scaling: loss ∝ compute^(-α) where α ≈ 0.05. Chinchilla refined this: for compute-optimal training, the ratio of parameters to tokens should be roughly 1:20. Chinchilla 70B trained on 1.4T tokens matched Gopher 280B trained on 300B tokens · confirming that most 2020-era models were under-trained. Post-Chinchilla, labs over-train aggressively (Llama 3 used 15T tokens on 70B params = 214:1 ratio) for better inference-time efficiency at the cost of training compute. #### Pretraining URL: https://benchgecko.ai/learn/glossary/pretraining Text reviewed 2026-10-05 TL;DR: The phase where a model learns from internet-scale text via next-token prediction · typically 15-30T tokens and millions of GPU hours. Data mix for 2026 pretraining: English web (~40%), code (~20%), books and scientific papers (~10%), multilingual (~20%), curated domain data (~10%). Quality filtering removes duplicate, toxic, or low-information content. Deduplication (MinHash, SimHash) is critical · duplicate data memorization is a concerning failure mode. Training compute measured in FLOPs: 10^25-10^26 for frontier class. FP8 pretraining has become standard since DeepSeek V3 (2024) demonstrated it at scale. See the related terms and live BenchGecko data for current examples. #### Embedding URL: https://benchgecko.ai/learn/glossary/embedding TL;DR: A dense numerical vector (256 to 8192 dimensions) that captures the semantic meaning of text, images, or audio. Modern embedding models are trained with contrastive learning: pairs of semantically related text (e.g., query and relevant doc) are pulled close in embedding space while random pairs are pushed apart. Dimensions range from 256 (fast) to 8192 (highest quality). Cosine similarity is the dominant distance metric. Good embedding models hit 70%+ MTEB leaderboard scores. Costs: $0.05-0.20/M tokens on the major APIs. Embedding a large corpus is a one-time cost; serving queries is fast and cheap. #### Multimodal URL: https://benchgecko.ai/learn/glossary/multimodal Text reviewed 2026-10-05 TL;DR: A model that handles multiple input or output types · text, image, audio, video · not just text alone. Two main architectures: (1) separate encoders per modality fused into a shared transformer (CLIP-style, earlier Gemini), and (2) native multimodal pretraining where tokens from different modalities share a unified embedding space (Gemini 1.5+, Chameleon). Native multimodal is more expressive but harder to train. Vision tokens use ViT (Vision Transformer) or patch-based encoders. Audio uses Whisper-style mel-spectrograms. Video extends images with temporal attention. Context windows for multimodal are often shorter · 1 minute of video can consume 10K+ tokens. #### Tool Use URL: https://benchgecko.ai/learn/glossary/tool-use TL;DR: When the model calls external functions (search, calculator, DB, API) during a response · the building block for agents. Flow: model generates a JSON tool call matching a schema → harness validates + executes the call → result is appended to the conversation → model continues. Parallel tool calls (invoke multiple at once) reduce latency. Tool descriptions matter enormously · a well-named and well-described tool is called reliably; a poorly-described one is ignored. Common tools: web search, calculator, database query, code execution, file operations, API calls to external services, MCP-exposed capabilities. Safety: destructive tools (file delete, send email) should require user approval. #### Prompt Engineering URL: https://benchgecko.ai/learn/glossary/prompt-engineering Text reviewed 2026-10-05 TL;DR: Crafting inputs to LLMs to get better outputs · few-shot examples, chain-of-thought triggers, role assignment, structured formats. Effective patterns: (1) Few-shot learning with 3-5 diverse examples, (2) Chain-of-thought triggers for math/logic, (3) Structured output via JSON schema or Pydantic, (4) Self-critique loops ("reflect on your answer"), (5) Decomposition ("break this into sub-tasks"). Anti-patterns: overly long preambles, conflicting instructions, emotional manipulation ("this is urgent!"). Claude's constitutional prompting, OpenAI's structured outputs, and Google's system-instruction API are the 2026 productized approaches. See the related terms and live BenchGecko data for current examples. #### Tokenizer URL: https://benchgecko.ai/learn/glossary/tokenizer Text reviewed 2026-10-05 TL;DR: The algorithm that converts raw text into tokens before the model processes them · BPE, SentencePiece, tiktoken. BPE (Byte Pair Encoding) builds vocabulary by greedily merging frequent byte pairs. SentencePiece works at raw byte level, language-agnostic. WordPiece (BERT-era) uses likelihood scoring. Modern frontier tokenizers have 200K+ vocabs to handle multilingual + code + mathematical notation efficiently. Larger vocab = fewer tokens per text but larger embedding tables. Tokenizer choice is locked at pretraining · swapping requires full retrain. #### Temperature URL: https://benchgecko.ai/learn/glossary/temperature TL;DR: A knob that controls how random or deterministic the AI's output is · 0 = same answer every time, higher = more varied. Math: logits are divided by temperature before softmax. Lower temperature sharpens the distribution; higher flattens it. Temperature 0 is not technically in the formula (would divide by zero) · it's implemented as argmax (pick highest logit). Common uses: 0 for structured extraction, factual Q&A; 0.3-0.7 for balanced chat; 0.8-1.2 for creative writing. Temperature and top-p are often used together but shouldn't both be tuned aggressively · pick one. #### Function Calling URL: https://benchgecko.ai/learn/glossary/function-calling Text reviewed 2026-10-05 TL;DR: The API mechanism that lets models invoke external tools by emitting structured JSON that matches a schema. Flow: (1) send request with function schemas, (2) model generates JSON tool call if it decides to use a tool, (3) your code validates + executes, (4) send result back as a tool message, (5) model continues or finishes. Parallel function calling (emit multiple tool calls per turn) reduces agent round-trips. OpenAI, Anthropic, Google, Mistral all support parallel calls. Quality of tool descriptions matters enormously · poorly described tools get skipped. #### AI Alignment URL: https://benchgecko.ai/learn/glossary/alignment TL;DR: The research discipline focused on making AI systems do what humans actually want · not just what they're told. Alignment research splits into (1) alignment techniques (how to train safe AI), and (2) safety research (what happens if techniques fail). Modern recipes: SFT → RLHF or DPO → safety tuning. Constitutional AI replaces human feedback with AI-generated critique following a written constitution. Scalable oversight (training AI to help evaluate AI) is an active research area. Evaluation: Anthropic's responsibility benchmarks, OpenAI's safety evals, red-team exercises. Alignment is necessary but not sufficient for safety; it is the foundation on which safety layers sit. #### Guardrails URL: https://benchgecko.ai/learn/glossary/guardrails Text reviewed 2026-10-05 TL;DR: Runtime safety filters around AI models · check inputs for attacks, check outputs for harms, block policy violations. Guardrail categories: (1) input filtering (prompt injection detection, jailbreak recognition, PII redaction), (2) output filtering (toxicity scoring, factual verification, format compliance), (3) rate/volume controls (anomaly detection, cost caps). Implementation: small classifier models, regex patterns, rule engines, or separate LLM calls dedicated to validation. Production systems stack multiple layers because no single approach catches everything. Trade-off: more guardrails = more latency and more false positives. #### Throughput URL: https://benchgecko.ai/learn/glossary/throughput TL;DR: The total tokens-per-second a serving cluster handles across ALL concurrent requests · vs tokens-per-second of a single request. Throughput depends on batching strategy, quantization, parallelism (tensor, pipeline, expert), and KV cache management. Continuous batching (vLLM, TRT-LLM) packs requests dynamically, drastically raising utilization. FP8 weights + activations doubles throughput on compatible hardware. Mixed-tier serving (small model for easy requests, big for hard) is an emerging pattern. Per-GPU throughput benchmark: H100 hits 2000-4000 tok/s aggregate on a 70B dense model with FP8 + continuous batching. #### Grounding URL: https://benchgecko.ai/learn/glossary/grounding TL;DR: Tethering AI output to verified source material via retrieval, citations, or tool calls · the counter to hallucination. Grounding spectrum: full RAG (retrieve + cite everything), partial RAG (retrieve for facts, LLM for synthesis), tool-grounded (model queries structured sources), and reasoning-grounded (model justifies each claim). Quality of grounding depends on retrieval recall and source authority. Enterprise AI increasingly demands citation-level grounding · "show me which document this came from" · for auditability and legal compliance. Perplexity popularized the citation-first UX pattern. #### Open weights URL: https://benchgecko.ai/learn/glossary/open-weights TL;DR: Open weights are model weights released for download, inspection, fine-tuning, or self-hosting under a license. Open weights means the model parameters are publicly available. It does not automatically mean open-source: the license may restrict commercial use, redistribution, or derivative models. For buyers and builders, open weights matter because they allow local deployment, privacy control, fine-tuning, and independence from one hosted API. #### FP8 URL: https://benchgecko.ai/learn/glossary/fp8 TL;DR: FP8 is an 8-bit floating point format used to speed up AI training and inference. FP8 reduces memory traffic and compute cost versus FP16 or BF16 while preserving enough numeric range for many transformer operations. It is common on newer accelerators and matters because lower precision can raise throughput, reduce cost, and change which hardware is competitive. #### FP16 URL: https://benchgecko.ai/learn/glossary/fp16 TL;DR: FP16 is a 16-bit floating point format widely used in neural network training and inference. FP16 stores numbers in half precision compared with FP32. It cuts memory and bandwidth use while keeping enough accuracy for most model workloads. FP16 remains a common baseline for comparing acceleration, quantization, VRAM use, and inference cost. #### Knowledge cutoff URL: https://benchgecko.ai/learn/glossary/knowledge-cutoff Text reviewed 2026-10-05 TL;DR: A knowledge cutoff is the point in time where a model's training data ends · the model cannot know about anything that happened later unless you give it that information. The stated cutoff and the effective cutoff can differ. BenchGecko measures the effective one in the Knowledge Horizon test: each model is asked about dated news events month by month, and the test finds where its answers stop being right. Retrieval (RAG) and web search tools are the usual way to give a model information newer than its cutoff. See live data: https://benchgecko.ai/gecko-tests/knowledge-horizon #### Test-time compute URL: https://benchgecko.ai/learn/glossary/test-time-compute Text reviewed 2026-10-05 TL;DR: Test-time compute is the computation a model spends at answer time · spending more (longer reasoning, several attempts) can raise accuracy without retraining the model. Common forms are long chain-of-thought reasoning, sampling several answers and voting, and search over candidate solutions with a verifier. The cost is paid on every request: more tokens, more latency and a higher bill, because reasoning tokens are billed as output. #### Model drift URL: https://benchgecko.ai/learn/glossary/model-drift Text reviewed 2026-10-05 TL;DR: Model drift is a change in a model's behavior while its name stays the same · caused by new weights, system prompts, safety filters or serving changes behind a stable API name. BenchGecko's Model Drift Index sends the same fixed probes to the same models every week and compares each run with the previous one: exact-answer accuracy, refusals and how the model names its maker. Small differences are expected because even temperature 0 is not perfectly deterministic, so drift is flagged only on large moves. See live data: https://benchgecko.ai/gecko-tests/model-drift-index #### LLM as a judge URL: https://benchgecko.ai/learn/glossary/llm-as-judge Text reviewed 2026-10-05 TL;DR: LLM as a judge means using one model to score or label the answers of another · fast and cheap evaluation for answers that code cannot grade exactly. Judges are cheaper and faster than human review but have known biases, such as preferring longer answers or answers in their own style. Good practice is a fixed rubric, a judge that sees only what it needs, and human spot checks. BenchGecko uses a judge model with a fixed rubric for the refusal labels in the Censorship Index and the Model Drift Index, and publishes every raw answer. #### Speculative decoding URL: https://benchgecko.ai/learn/glossary/speculative-decoding Text reviewed 2026-10-05 TL;DR: Speculative decoding speeds up generation without changing the output · a cheap draft model guesses several tokens and the main model checks them all at once. With the standard acceptance rule, the output follows the same distribution as the large model alone, so quality is unchanged. Speed-ups depend on how often the draft is right, which is higher for predictable text such as code. ### Pricing #### Cache Hit Rate URL: https://benchgecko.ai/learn/glossary/cache-hit-rate Text reviewed 2026-10-05 TL;DR: The percentage of input tokens served from a provider's prompt cache · modern providers charge 10% of list price on cache hits, so hit rate directly sets effective input cost. Anthropic's caching has a 5-minute or 1-hour TTL · you pay a premium to write to cache (1.25× list for 5-min, 2× for 1-hour), then 0.1× for reads. OpenAI caches automatically with shorter TTL and no write premium. Google Gemini cache is context-size dependent. Hit rate depends on prompt design: keep static content at the front, split cacheable blocks from dynamic ones, and reuse the same prefix across sessions. Most production AI apps see 60-95% cache hit rates with proper design. #### Batched Inference URL: https://benchgecko.ai/learn/glossary/batched-inference Text reviewed 2026-10-05 TL;DR: Batch APIs let you submit thousands of prompts for offline processing at 50% of synchronous pricing · turnaround up to 24 hours. Batch throughput bypasses rate limits · you can submit 50K prompts even if your sync rate limit is 1000 RPM. The 50% discount applies to both input and output tokens. Providers run batches on leftover capacity, which is why the SLA is "up to 24 hours." Typical real-world turnaround: 1-4 hours. Batch pricing combined with prompt caching gives an effective 10-20× cost reduction vs naive synchronous calls. #### Reserved Capacity URL: https://benchgecko.ai/learn/glossary/reserved-capacity Text reviewed 2026-10-05 TL;DR: Reserved capacity sells dedicated inference slots at a flat hourly rate · you buy tokens-per-second, not per-call pricing. Reserved capacity breakeven depends on utilization. Azure OpenAI PTUs are similar: reserve 300-500 TPS blocks. Google offers minute-granularity on Vertex. Discounts increase with commitment length · 1-year reservations are 30-50% cheaper than on-demand. Current prices for every tracked model are on the BenchGecko pricing pages. #### Per-Request Pricing URL: https://benchgecko.ai/learn/glossary/per-request-pricing TL;DR: A flat fee per API request regardless of input or output size · common for image generation, web search, and some premium agent products. Per-request pricing tradeoff: predictability vs fairness. For fixed-output features (image generation, fixed-size embeddings, tool calls), flat fee is simple. For variable-length outputs, per-token is fairer. Some providers blend · flat fee plus per-token for output length. Use cases best served by per-request: web search, image gen, function call billing, tool invocations. Enterprise customers often prefer per-request for predictable budgeting. #### Tiered Pricing URL: https://benchgecko.ai/learn/glossary/tiered-pricing Text reviewed 2026-10-05 TL;DR: Tiered pricing drops per-unit cost as volume crosses defined thresholds · common in AI pricing above enterprise tier. Tiered pricing differs from volume discount contracts. Tiered is automatic · you cross a threshold, the new rate kicks in mid-billing-period. Volume discount is a pre-negotiated flat rate applied to all consumption (simpler for finance). Most providers offer both shapes. Tiered benefits customers whose volume is uncertain · you pay less without negotiating. The trap: tier boundaries can create cliff effects where running 10% more volume saves 20% on total bill. #### Volume Discounts URL: https://benchgecko.ai/learn/glossary/volume-discounts Text reviewed 2026-10-05 TL;DR: Volume discounts replace public per-token pricing with a pre-negotiated flat rate · typical shape of enterprise AI contracts. Volume discount contracts include clauses around usage: minimum commits, rollover rules, overage rates, and term length. Rollover is negotiated · some contracts allow 3-month rollover of unused tokens. Volume discount contracts cover multiple models · if public rates drop 50% mid-contract, most contracts ratchet to best-available price. Current prices for every tracked model are on the BenchGecko pricing pages. #### Multi-Modal Pricing URL: https://benchgecko.ai/learn/glossary/multi-modal-pricing Text reviewed 2026-10-05 TL;DR: Multi-modal models charge different rates for different input types · images, audio, and video have their own per-unit prices alongside text tokens. Image tokens: providers convert images to token equivalents via tile-based encoding. OpenAI: ~170 tokens per 512×512 tile. Anthropic: 1,200 tokens per image (approximate). Video: Gemini counts 258 tokens per frame at 1fps; you can adjust fps. Output tokens are almost always text-only and priced as regular output. Tracking multi-modal consumption requires per-input metering. Current prices for every tracked model are on the BenchGecko pricing pages. #### Reasoning Token Billing URL: https://benchgecko.ai/learn/glossary/reasoning-token-billing Text reviewed 2026-10-05 TL;DR: Reasoning models (o-series, Claude extended thinking, DeepSeek R1) bill hidden thinking tokens at the output rate · total cost can be 2-10× a non-reasoning query. Reasoning token pricing forces app builders to cap reasoning budgets. OpenAI added a `max_output_tokens` parameter that hard-caps total tokens; most APIs added `reasoning_effort` (low/medium/high) to tune budget. Typical reasoning-to-answer ratios: math-heavy tasks 20:1, simple Q&A 2:1. For high-traffic apps, rolling reasoning into an initial cheap pass and falling back to a reasoning model on hard cases is now standard practice. #### Function Call Billing URL: https://benchgecko.ai/learn/glossary/function-call-billing Text reviewed 2026-10-05 TL;DR: When a model emits a function call, the JSON tool-invocation counts as output tokens · cheap tools still carry token cost, and chained tool calls compound fast. OpenAI now supports "parallel tool calls" where one assistant turn emits multiple tool invocations in a single response · reduces per-call overhead. MCP servers add tool definitions via the clients' session context · those definitions hit input tokens once per new session. Caching tool definitions (Anthropic cache_control) brings input cost down to near-zero on subsequent calls. Current prices for every tracked model are on the BenchGecko pricing pages. #### Spot Pricing URL: https://benchgecko.ai/learn/glossary/spot-pricing Text reviewed 2026-10-05 TL;DR: Spot pricing sells preemptible GPU capacity at 60-90% off on-demand · you can be evicted with 2 minutes notice but prices are dramatically lower. Spot availability fluctuates · when a cloud provider's on-demand demand spikes, spot instances are reclaimed. Eviction notice is typically 30 seconds to 2 minutes. Training frameworks (PyTorch Lightning, DeepSpeed, HF Accelerate) support spot-friendly checkpointing. For inference, spot is used for batch endpoints that can tolerate partial completion. Runpod, Vast.ai, and Lambda Cloud run spot-style markets · often cheaper than hyperscaler spot. #### BYOK (BYOK) URL: https://benchgecko.ai/learn/glossary/byok Text reviewed 2026-10-05 TL;DR: A pricing model where you provide your own OpenAI/Anthropic/Google API key, and the app uses it to make calls on your behalf. At high volume, BYOK is cheaper because you bypass the app's margin. Breakeven point is typically 500K-2M tokens/month. BYOK is common for enterprise deployments (security, billing consolidation) and power users. Some apps run hybrid · subscription covers a baseline, BYOK for overage. Current prices for every tracked model are on the BenchGecko pricing pages. #### Input Tokens URL: https://benchgecko.ai/learn/glossary/input-tokens Text reviewed 2026-10-05 TL;DR: The tokens in your prompt · billed per million, typically 3-5× cheaper than output tokens. Input token cost scales with prompt length. Long system prompts + full retrieval + long conversation history = expensive. Most production apps land at 70-85% of total token cost on input despite paying more per output token, because inputs are typically 5-20× longer than outputs. Current prices for every tracked model are on the BenchGecko pricing pages. #### Output Tokens URL: https://benchgecko.ai/learn/glossary/output-tokens Text reviewed 2026-10-05 TL;DR: The tokens the model writes back to you · priced 3-5× higher than input because decoding is sequential. Output tokens are expensive because they're generated one at a time during the decode phase. Each token requires a full pass through the model reading the KV cache · memory bandwidth bound. This is why faster GPUs and newer memory generations matter so much for inference economics. Speculative decoding yields 2-3× throughput on output tokens. FP8 and INT8 quantization at both weights and activations doubles output throughput again. #### Cache Pricing URL: https://benchgecko.ai/learn/glossary/cache-pricing Text reviewed 2026-10-05 TL;DR: A provider feature that gives you discounted input pricing for prompts the system has already processed · 50-90% savings. Cache pricing maps onto the prefill/decode split. The prefill work for a stable prefix (system prompt, tools) can be amortized across many requests. Providers expose this as a pricing tier · mark which prefix portion is stable, and subsequent requests hit cache at discounted rates. Caches expire after minutes of idle (varies by provider). Multi-turn chat workloads with stable system prompts see 60-80% total cost reduction when cache-aware. #### Free Tier URL: https://benchgecko.ai/learn/glossary/free-tier Text reviewed 2026-10-05 TL;DR: The no-cost tier for AI APIs · usually rate-limited, quota-capped, or required-to-include-attribution. Free tier models range: Groq free (limited to 30 req/min), Gemini free (15 req/min, token quotas), Together free $5 credit, Anthropic $5 signup credit, OpenAI $5 signup credit. Sustainable free tier: self-hosted open-weight models (Llama, Qwen, Mistral) on your own GPU. BYOK in apps like Cline + your own Anthropic/OpenAI key is "free app, pay tokens" model. Tracking all this is a headache · see /pricing/free for the live sheet. #### Arbitrage (AI Pricing) URL: https://benchgecko.ai/learn/glossary/arbitrage Text reviewed 2026-10-05 TL;DR: Current prices for every tracked model are on the BenchGecko pricing pages. Arbitrage dimensions: token price, latency, throughput, reliability, compliance. Cheapest provider often has worst latency; fastest often charges premium. OpenRouter aggregates 200+ providers and routes automatically. Self-hosting hits zero marginal cost but adds engineering overhead. BenchGecko's /pricing/arbitrage/[model-slug] tracks every hosted instance with price + latency per provider. #### Tokenizer tax URL: https://benchgecko.ai/learn/glossary/tokenizer-tax Text reviewed 2026-10-05 TL;DR: The tokenizer tax is the extra cost of non-English text · tokenizers split many languages into more tokens than English, and models bill per token. BenchGecko's Tokenizer Tax test sends the same text in several languages to each model and records how many input tokens the provider counts, relative to English. Differences between models come from their tokenizers: larger multilingual vocabularies usually lower the tax. See live data: https://benchgecko.ai/gecko-tests/tokenizer-tax #### Batch API URL: https://benchgecko.ai/learn/glossary/batch-api Text reviewed 2026-10-05 TL;DR: A batch API takes a file of requests and returns the results later, usually within a day, at a discount to real-time pricing. BenchGecko lists batch pricing as separate model entries marked "(batch)", so you can compare them directly with real-time prices. The trade-off is latency: results arrive within the provider's batch window, not in seconds. See live data: https://benchgecko.ai/pricing ### Economy #### Revenue per Employee URL: https://benchgecko.ai/learn/glossary/revenue-per-employee Text reviewed 2026-10-05 TL;DR: Revenue per employee measures operating leverage · AI labs hit $1-5M/employee while traditional SaaS ranges $200-400K. RPE is a late-stage efficiency signal. Early stage, RPE is meaningless (too few humans, too little revenue). AI labs trend high because compute + data do most of the work · engineers build and maintain systems that scale without linear headcount growth. This is also why AI labs command higher valuations per employee than SaaS. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Gross Margin (AI Labs) URL: https://benchgecko.ai/learn/glossary/gross-margin-ai Text reviewed 2026-10-05 TL;DR: AI lab gross margins (50-70%) sit below pure SaaS (80-90%) because inference compute is a variable cost, not amortized software. Training cost is a separate line (R&D) and doesn't hit COGS directly. Efficient architectures (MoE, distillation) and cheaper GPUs (Blackwell, H200, MI300) push COGS down year over year · AI lab gross margins are trending up as compute gets cheaper. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Burn Multiple URL: https://benchgecko.ai/learn/glossary/burn-multiple Text reviewed 2026-10-05 TL;DR: Burn multiple = net burn / net new ARR · sub-1× is great · AI startups regularly run 3-5× due to compute cost. AI burn multiples decompose into: R&D (model training, salaries), inference COGS, sales, and GTM spend. Inference COGS grows with usage but has less up-front cost. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Customer Concentration URL: https://benchgecko.ai/learn/glossary/customer-concentration Text reviewed 2026-10-05 TL;DR: Customer concentration measures how much revenue depends on a few customers · AI labs often sit at 40-60% from top 10 vs SaaS at <20%. High concentration = fragile (losing one customer hurts) but also validating (enterprise buyers staked their AI strategy). Low concentration = robust but usually means lower absolute revenue. Current figures, with dates and sources, are on the BenchGecko economy pages. #### GMV (GMV) URL: https://benchgecko.ai/learn/glossary/gmv Text reviewed 2026-10-05 TL;DR: Gross Merchandise Value measures total transaction value flowing through a platform · relevant for AI marketplaces (model routers, MCP registries, agent marketplaces). GMV matters for marketplaces because network effects scale with transaction volume, not with platform take rate. OpenRouter's revenue is <5% of its GMV but the GMV flows determine which providers care to integrate. This is why AI marketplace valuations can look inflated on pure revenue. Current figures, with dates and sources, are on the BenchGecko economy pages. #### LTV / CAC URL: https://benchgecko.ai/learn/glossary/ltv-cac Text reviewed 2026-10-05 TL;DR: LTV/CAC compares lifetime revenue per customer to acquisition cost · >3× is the SaaS healthy benchmark · AI consumer apps are 2-4× so far. LTV calculation is fragile · typical formula: (average monthly revenue × gross margin) / monthly churn rate. If CAC is $80, LTV/CAC = 3×. Higher-churn apps (weekly AI image generators) run lower. Enterprise AI is different · LTV extends 3-5 years with NRR >100%, CAC is heavy but ratio often 5-10×. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Investor Power Map URL: https://benchgecko.ai/learn/glossary/investor-power-map Text reviewed 2026-10-05 TL;DR: The investor power map tracks which funds + corporate strategics dominate each tier of AI deals · concentration at the top matters for terms. Power dynamics: corporate strategics (Microsoft + OpenAI, Amazon + Anthropic) shape not just capital but also compute access and distribution. Thrive Capital + Josh Kushner and a16z both repeatedly invest in frontier labs and agent companies, creating network effects. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Dilution URL: https://benchgecko.ai/learn/glossary/dilution Text reviewed 2026-10-05 TL;DR: Dilution is the % ownership reduction existing shareholders experience after a new round · typical 15-25% per standard round, more at hot pre-seed. Dilution math: new ownership = existing shares / (existing + new shares). Pre-money valuation × ownership-pre = post-money × ownership-post. Pro-rata rights let existing investors buy in new rounds to maintain %. Option pool top-ups also dilute · companies usually top up to 10-15% post-Series B. Founder dilution trajectory: from 100% at founding to 15-30% post-Series C is typical. Frontier AI lab founders often below 10% post multiple mega-rounds. #### ESOP (ESOP) URL: https://benchgecko.ai/learn/glossary/esop Text reviewed 2026-10-05 TL;DR: ESOP is the employee stock pool · typically 10-20% of cap table · expanded each round to attract talent. ESOP mechanics: company board approves share pool size. Shares are granted to employees at fair market value (409A) on grant date with a vesting schedule (usually 4-year, 1-year cliff). Unused pool rolls forward; depleted pool triggers top-up. Common pattern: 1% ESOP for senior engineers, 0.1-0.5% for mid-level, 0.01-0.05% for junior. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Ratchet URL: https://benchgecko.ai/learn/glossary/ratchet TL;DR: A ratchet adjusts preferred-share conversion price if a later round is at a lower valuation · full ratchet is harsh, weighted-average is standard. Full ratchet is founder-hostile · it can wipe common equity to near zero in a severe down-round. Weighted-average is milder. Most Series B+ AI deals use weighted-average. The "broad-based" weighted-average is the market default. Some hot AI rounds (Mistral, Anthropic early) included pay-to-play clauses · existing investors must participate in follow-ons to retain anti-dilution protections. #### ARR (ARR) URL: https://benchgecko.ai/learn/glossary/arr Text reviewed 2026-10-05 TL;DR: Annualized Recurring Revenue · take current monthly recurring revenue × 12. The standard AI/SaaS revenue metric. ARR methodology varies: some companies include one-time revenue (contract bookings ÷ contract length), some include usage-based. Pure SaaS ARR = MRR × 12 from recurring contracts only. Most AI companies include usage-based revenue because pure subscription misses the heavy API users. Current figures, with dates and sources, are on the BenchGecko economy pages. #### AI Capex URL: https://benchgecko.ai/learn/glossary/capex Text reviewed 2026-10-05 TL;DR: The billions hyperscalers and AI labs spend each year on GPUs, datacenters, and training clusters · the #1 driver of AI spending narrative. Capex accounting: GPUs + data center shell + power infrastructure + networking + cooling. Data center shell + power can be another 30% of total spend. The capex-to-revenue ratio across AI hyperscalers exceeded 50% in 2025 · historically extreme. This gap drives the AI Bubble Index component on BenchGecko. Current figures, with dates and sources, are on the BenchGecko economy pages. #### P/S Ratio (P/S) URL: https://benchgecko.ai/learn/glossary/ps-ratio Text reviewed 2026-10-05 TL;DR: Valuation divided by annual revenue · the single number that captures whether an AI company is fairly priced, frothy, or bubbled. P/S ratio benchmarks: healthy SaaS 5-15×, hypergrowth SaaS 20-50×, frothy 50-100×, bubble 100+×. Drivers: hyperscaler AI revenue counted as "software" growth, investor FOMO, and genuine belief in 10× productivity impact. Outlier cases (Cursor 200×, Perplexity 100×) reflect pricing power premium over incumbent AI. #### AI Bubble Index URL: https://benchgecko.ai/learn/glossary/bubble-index Text reviewed 2026-10-05 TL;DR: A 0-1000%+ composite score measuring how bubbled (or not) the AI sector is · BenchGecko's flagship economy indicator. Computed daily from BenchGecko tracked companies. Scope checkboxes let users define the index (Pure AI only? Include Big Tech? Chips? Foundries?) · each scope produces a different reading. Pure AI scope typically reads 400-500% · overheated. The methodology is published openly on BenchGecko · not a black box. Current figures, with dates and sources, are on the BenchGecko economy pages. #### Burn Rate URL: https://benchgecko.ai/learn/glossary/burn-rate Text reviewed 2026-10-05 TL;DR: How fast an AI company is spending investor cash · usually measured as monthly cash outflow minus revenue. AI burn rate composition: compute (40-60%), salaries (25-40%), data acquisition (5-15%), marketing (5-20%). Frontier labs inverted from SaaS burn profiles · compute is the dominant line item. Smaller AI startups often can't out-compete frontier on quality, so they differentiate on vertical (Cursor on coding, Character on companionship) to contain burn. Current figures, with dates and sources, are on the BenchGecko economy pages. ### Agents #### Devin URL: https://benchgecko.ai/learn/glossary/devin Text reviewed 2026-10-05 TL;DR: Devin is Cognition AI's autonomous SWE agent · first to claim end-to-end ticket-to-PR automation with browser, shell, and editor access. Devin pioneered the "agent as cloud employee" pattern. Each Devin session runs in an isolated VM with persistent state · filesystem, browser history, installed packages · so multi-hour tasks survive across runs. Integration with Slack, Jira, GitHub, Linear turned it into the model for later agents (Cursor Background, OpenAI Codex). See the agents and MCP pages on BenchGecko for current tools. See live data: https://benchgecko.ai/agent/devin #### Replit Agent URL: https://benchgecko.ai/learn/glossary/replit-agent Text reviewed 2026-10-05 TL;DR: Replit's browser-native coding agent · takes a natural-language prompt and ships a running app, database, and deploy URL in one flow. The harness has access to Replit's container runtime, PostgreSQL instances, object storage, and secrets manager · all provisioned on demand. The Nix-based environment lets it install any language runtime without user config. See the agents and MCP pages on BenchGecko for current tools. See live data: https://benchgecko.ai/agents #### Operator URL: https://benchgecko.ai/learn/glossary/operator Text reviewed 2026-10-05 TL;DR: OpenAI Operator is a computer-use agent · controls a cloud browser to book, shop, research, and fill forms on behalf of the user. Operator's model (internally called CUA) is fine-tuned from GPT-4 on screen-capture + action-label pairs. It receives a screenshot, predicts a coordinate click or keystroke, and iterates. Human handoff triggers on login, CAPTCHAs, and payments (Operator never enters credit cards autonomously). Benchmarks: 38.1% on WebArena, 87% on WebVoyager at launch. Pricing is bundled into Pro tier. See live data: https://benchgecko.ai/agents #### Manus URL: https://benchgecko.ai/learn/glossary/manus Text reviewed 2026-10-05 TL;DR: Manus is a China-based autonomous agent that went viral March 2025 for booking flights, researching markets, and filing SEC documents end-to-end · invitation-only launch. Manus uses Claude 3.7 Sonnet as its primary reasoning model (public reverse-engineering revealed the API calls). The harness is similar to Devin · cloud VM with persistent state · but specializes in multi-hour research tasks producing long-form deliverables. Notable for shipping multi-modal input (image + PDF + audio) and long-horizon reasoning that makes other agents feel constrained. See the agents and MCP pages on BenchGecko for current tools. See live data: https://benchgecko.ai/agents #### Cline URL: https://benchgecko.ai/learn/glossary/cline Text reviewed 2026-10-05 TL;DR: Cline is the original open-source VSCode coding agent · users bring their own API key, the extension does the work. Cline's harness is a tool-use loop: read_file, write_file, execute_command, browser_action, ask_user. Every tool call requires user approval by default (toggleable to auto-mode). Context management includes automatic file-change diffs, directory trees, and open-file tracking. The open-source codebase forked into Roo Code (automation-heavy) and influenced Continue, Aider, and dozens of other agents. Apache-2.0 licensed · no commercial restrictions. See live data: https://benchgecko.ai/agents #### Aider URL: https://benchgecko.ai/learn/glossary/aider TL;DR: Aider is a terminal-native AI pair programmer with first-class git integration · every change lands as a clean commit. Aider's harness is minimalist: no GUI, no MCP, no browser · just edit-run-commit. Under the hood, it maintains a "repo map" (symbol graph) for context, uses SEARCH/REPLACE diffs to minimize token cost, and auto-tests after each change if a test command is configured. The Aider Polyglot benchmark (Rust, Python, JS, Go, C++, Java, Swift + more) is widely cited · every major model now reports its Aider score. See live data: https://benchgecko.ai/benchmark/aider-polyglot #### Continue.dev URL: https://benchgecko.ai/learn/glossary/continue-dev Text reviewed 2026-10-05 TL;DR: Continue is an open-source coding assistant that runs inside VSCode and JetBrains · inline completions, chat, and local-model support. Continue ships two primary models: a fast autocomplete model (small, local-friendly) and a chat model (larger, remote). The `config.json` system lets teams standardize on specific models and prompts across developers. Enterprise features include on-prem deployment, audit logs, and SOC2 compliance. See the agents and MCP pages on BenchGecko for current tools. See live data: https://benchgecko.ai/agents #### Codex CLI URL: https://benchgecko.ai/learn/glossary/codex-cli Text reviewed 2026-10-05 TL;DR: Codex CLI is OpenAI's open-source terminal agent · GPT-powered, repo-aware, installed with npm. Codex CLI is Apache-2.0 licensed · same repo as the Codex IDE (VSCode) and Codex Cloud (web SaaS). The harness exposes shell commands, file diffs, and HTTP calls. Token cost is high · a single task may burn $2-$5. See live data: https://benchgecko.ai/agents #### Roo Code URL: https://benchgecko.ai/learn/glossary/roo-code Text reviewed 2026-10-05 TL;DR: Roo Code is a Cline fork that adds auto-execution modes, multi-tab orchestration, and an Orchestrator-Worker agent pattern. The orchestrator-worker pattern lets users define a parent agent that delegates subtasks to specialist child agents. Example: orchestrator plans a migration, worker agents handle backend, frontend, and tests in parallel. Each subtask has its own approval flow. Roo Code also supports custom "modes" · named agent personas with specific prompts, tool subsets, and auto-approve rules. Users can ship modes as shareable files. See live data: https://benchgecko.ai/agents #### Windsurf URL: https://benchgecko.ai/learn/glossary/windsurf-agent Text reviewed 2026-10-05 TL;DR: Windsurf is Codeium's VSCode-fork IDE with the Cascade multi-file agent · acquired by Cognition AI in 2025 to pair with Devin. Cascade pioneered "flow awareness" · the agent watches your recent actions (file opens, edits, terminal runs) and adjusts its plan without explicit prompting. This was distinct from Cursor's more reactive model. Windsurf retains local IDE UX while Devin lives in the cloud · the pair positions Cognition as both local + remote agent provider. See live data: https://benchgecko.ai/agents #### Zed Agent URL: https://benchgecko.ai/learn/glossary/zed-agent TL;DR: Zed Agent is the AI assistant built into the Zed editor · a Rust-native, collaborative code editor from the Atom creators. Zed is open source (GPL-3). The agent supports Anthropic, OpenAI, OpenRouter, and Ollama. Inline completions use a small local model; chat and edits use remote. MCP server support was shipped in 2025 · Zed is one of the reference MCP clients. Zed's differentiation from Cursor/Windsurf is performance · the editor itself is faster, and the AI integration respects that (streaming, incremental diffs). See live data: https://benchgecko.ai/agents #### Tabby URL: https://benchgecko.ai/learn/glossary/tabby TL;DR: Tabby is a self-hosted AI coding assistant · Copilot-shaped autocomplete you run on your own hardware for privacy. Tabby ships pre-configured with StarCoder2 and other open models. Admins pick the model based on GPU budget · a single A100 serves ~10-20 concurrent developers. Features beyond autocomplete include chat, code-search, and a team dashboard with analytics. Enterprise tier adds SSO, audit logs, and on-prem LDAP. Apache-2.0 with a paid enterprise offering. See live data: https://benchgecko.ai/agents #### Cody URL: https://benchgecko.ai/learn/glossary/cody Text reviewed 2026-10-05 TL;DR: Cody is Sourcegraph's enterprise AI coding assistant · leverages their code-graph index for deep repo-wide context. Cody's context engine combines Sourcegraph Code Search, embeddings, and graph traversal. When asked "how is this function used?" Cody can answer with actual call sites, not hallucinations. Model-pluggable · Claude, GPT, Mixtral, StarCoder. Enterprise deployment supports on-prem with BYO-API-key or BYO-model. SOC2 + HIPAA certified. See live data: https://benchgecko.ai/agents #### Goose URL: https://benchgecko.ai/learn/glossary/goose TL;DR: Goose is an open-source local coding agent from Block (Square) · MCP-first, works offline with local models, published Apache-2.0. Goose's harness is minimal: a planning loop, tool invocation, and response streaming. MCP support is the core feature · install servers for GitHub, Jira, shell, filesystem, and Goose exposes them as a unified toolset. The desktop app (Tauri-based) adds a chat UI and session persistence. Developers at Block use Goose with internal MCP servers that expose billing data, deploy systems, and codebases. See live data: https://benchgecko.ai/agents #### Claude Code URL: https://benchgecko.ai/learn/glossary/claude-code TL;DR: Claude Code is a command-line coding agent built by Anthropic that lives in your terminal and edits files, runs commands, and searches your codebase with direct Claude model access. Claude Code is a command-line coding agent built by Anthropic that lives in your terminal and edits files, runs commands, and searches your codebase with direct Claude model access. Built by Anthropic · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Claude Code's GitHub stars, SWE-bench scores, and adoption over time on /agent/claude-code. See live data: https://benchgecko.ai/agent/claude-code #### Cursor URL: https://benchgecko.ai/learn/glossary/cursor TL;DR: Cursor is a fork of VS Code rebuilt around an AI pair programmer. Cursor is a fork of VS Code rebuilt around an AI pair programmer. Includes inline tab-completion, agent mode (Composer), and multi-file edit flows powered by frontier models. Built by Anysphere · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Cursor's GitHub stars, SWE-bench scores, and adoption over time on /agent/cursor. See live data: https://benchgecko.ai/agent/cursor #### Windsurf URL: https://benchgecko.ai/learn/glossary/windsurf TL;DR: Windsurf is Codeium's purpose-built agentic IDE with Cascade, a flow-state collaborative agent that keeps context across your whole project. Windsurf is Codeium's purpose-built agentic IDE with Cascade, a flow-state collaborative agent that keeps context across your whole project. Built by Codeium · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Windsurf's GitHub stars, SWE-bench scores, and adoption over time on /agent/windsurf. See live data: https://benchgecko.ai/agent/windsurf #### OpenHands URL: https://benchgecko.ai/learn/glossary/openhands TL;DR: OpenHands (formerly OpenDevin) is a community-built autonomous software development agent that matches frontier closed-source agents on SWE-bench Verified. OpenHands (formerly OpenDevin) is a community-built autonomous software development agent that matches frontier closed-source agents on SWE-bench Verified. Runs any model via LiteLLM. Built by All Hands AI · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks OpenHands's GitHub stars, SWE-bench scores, and adoption over time on /agent/openhands. See live data: https://benchgecko.ai/agent/openhands #### Continue URL: https://benchgecko.ai/learn/glossary/continue TL;DR: Continue is an open-source IDE extension for building your own AI code assistant with custom models, context providers, and slash commands. Continue is an open-source IDE extension for building your own AI code assistant with custom models, context providers, and slash commands. Works in VS Code and JetBrains. Built by Continue · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Continue's GitHub stars, SWE-bench scores, and adoption over time on /agent/continue. See live data: https://benchgecko.ai/agent/continue #### GitHub Copilot URL: https://benchgecko.ai/learn/glossary/github-copilot TL;DR: GitHub Copilot is the AI pair programmer from GitHub and Microsoft. GitHub Copilot is the AI pair programmer from GitHub and Microsoft. Inline code completions, chat, and agent mode across every major IDE and pull-request workflow. Built by GitHub · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks GitHub Copilot's GitHub stars, SWE-bench scores, and adoption over time on /agent/github-copilot. See live data: https://benchgecko.ai/agent/github-copilot #### Jules URL: https://benchgecko.ai/learn/glossary/jules TL;DR: Jules is Google's asynchronous coding agent. Jules is Google's asynchronous coding agent. Clones your repo, plans a fix, runs code, and opens a pull request · all in a Cloud VM running Gemini. Built by Google · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks Jules's GitHub stars, SWE-bench scores, and adoption over time on /agent/jules. See live data: https://benchgecko.ai/agent/jules #### Sweep URL: https://benchgecko.ai/learn/glossary/sweep TL;DR: Sweep is an AI junior developer that turns bug reports and feature requests into code changes as pull requests, running entirely from GitHub issues. Sweep is an AI junior developer that turns bug reports and feature requests into code changes as pull requests, running entirely from GitHub issues. Built by Sweep AI · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks Sweep's GitHub stars, SWE-bench scores, and adoption over time on /agent/sweep. See live data: https://benchgecko.ai/agent/sweep #### GPT Engineer URL: https://benchgecko.ai/learn/glossary/gpt-engineer TL;DR: GPT Engineer generates an entire codebase from a natural language spec. GPT Engineer generates an entire codebase from a natural language spec. One of the early viral autonomous coding agents. Now maintained by a community. Built by Anton Osika · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks GPT Engineer's GitHub stars, SWE-bench scores, and adoption over time on /agent/gpt-engineer. See live data: https://benchgecko.ai/agent/gpt-engineer #### AutoGPT URL: https://benchgecko.ai/learn/glossary/autogpt TL;DR: AutoGPT is the open-source autonomous agent that kicked off the entire category. AutoGPT is the open-source autonomous agent that kicked off the entire category. Now a low-code platform for building, running, and sharing continuous agents. Built by Significant Gravitas · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks AutoGPT's GitHub stars, SWE-bench scores, and adoption over time on /agent/autogpt. See live data: https://benchgecko.ai/agent/autogpt #### MetaGPT URL: https://benchgecko.ai/learn/glossary/metagpt TL;DR: MetaGPT models a software company with product managers, architects, engineers and QAs. MetaGPT models a software company with product managers, architects, engineers and QAs. One natural-language prompt produces user stories, APIs, and working code through role-based multi-agent collaboration. Built by DeepWisdom · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks MetaGPT's GitHub stars, SWE-bench scores, and adoption over time on /agent/metagpt. See live data: https://benchgecko.ai/agent/metagpt #### SWE-agent URL: https://benchgecko.ai/learn/glossary/swe-agent TL;DR: SWE-agent is the research agent from the Princeton NLP group behind SWE-bench. SWE-agent is the research agent from the Princeton NLP group behind SWE-bench. Introduced the agent-computer interface (ACI) idea · a restricted terminal-like API optimized for LLM tool use. Built by Princeton NLP · proprietary · category: autonomous. Supported models: multiple. BenchGecko tracks SWE-agent's GitHub stars, SWE-bench scores, and adoption over time on /agent/swe-agent. See live data: https://benchgecko.ai/agent/swe-agent #### Plandex URL: https://benchgecko.ai/learn/glossary/plandex TL;DR: Plandex is an open-source terminal AI coding agent optimized for large, multi-file tasks. Plandex is an open-source terminal AI coding agent optimized for large, multi-file tasks. Uses a separate context management layer to handle projects that exceed model context windows. Built by Plandex · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Plandex's GitHub stars, SWE-bench scores, and adoption over time on /agent/plandex. See live data: https://benchgecko.ai/agent/plandex #### Zed URL: https://benchgecko.ai/learn/glossary/zed TL;DR: Zed is a Rust-built code editor from the Atom team. Zed is a Rust-built code editor from the Atom team. Features real-time collaboration, GPU-accelerated rendering, and a built-in AI assistant with agentic edit prediction. Built by Zed Industries · proprietary · category: assistive. Supported models: multiple. BenchGecko tracks Zed's GitHub stars, SWE-bench scores, and adoption over time on /agent/zed. See live data: https://benchgecko.ai/agent/zed #### AutoGen URL: https://benchgecko.ai/learn/glossary/autogen TL;DR: AutoGen is Microsoft's open-source framework for building multi-agent AI applications. AutoGen is Microsoft's open-source framework for building multi-agent AI applications. Agents converse, delegate sub-tasks, and collaborate to solve complex problems. Built by Microsoft · proprietary · category: infrastructure. Supported models: multiple. BenchGecko tracks AutoGen's GitHub stars, SWE-bench scores, and adoption over time on /agent/autogen. See live data: https://benchgecko.ai/agent/autogen #### CrewAI URL: https://benchgecko.ai/learn/glossary/crewai TL;DR: CrewAI is a framework for orchestrating autonomous AI agents with roles, goals, and tools. CrewAI is a framework for orchestrating autonomous AI agents with roles, goals, and tools. Build crews of specialized agents that collaborate on tasks · used by over 60% of the Fortune 500. Built by crewAIInc · proprietary · category: infrastructure. Supported models: multiple. BenchGecko tracks CrewAI's GitHub stars, SWE-bench scores, and adoption over time on /agent/crewai. See live data: https://benchgecko.ai/agent/crewai #### LangGraph URL: https://benchgecko.ai/learn/glossary/langgraph TL;DR: LangGraph is LangChain's framework for building stateful, multi-actor agents as graphs. LangGraph is LangChain's framework for building stateful, multi-actor agents as graphs. Models agent flows as nodes and edges with full control over state, branching, and human-in-the-loop. Built by LangChain · proprietary · category: infrastructure. Supported models: multiple. BenchGecko tracks LangGraph's GitHub stars, SWE-bench scores, and adoption over time on /agent/langgraph. See live data: https://benchgecko.ai/agent/langgraph #### Smolagents URL: https://benchgecko.ai/learn/glossary/smolagents TL;DR: Smolagents is Hugging Face's minimalist agent library. Smolagents is Hugging Face's minimalist agent library. First-class support for code-writing agents where the model writes Python directly as its action space instead of JSON tool calls. Built by Hugging Face · proprietary · category: infrastructure. Supported models: multiple. BenchGecko tracks Smolagents's GitHub stars, SWE-bench scores, and adoption over time on /agent/smolagents. See live data: https://benchgecko.ai/agent/smolagents #### Computer use URL: https://benchgecko.ai/learn/glossary/computer-use Text reviewed 2026-10-05 TL;DR: Computer use lets a model control a desktop or browser like a person · it sees screenshots and sends mouse and keyboard actions. Computer use is slower and less reliable than API calls because every step goes through vision and many actions. Benchmarks such as OSWorld measure success rates on real desktop tasks. Sandboxing matters: an agent with a real screen can click anything a user could. #### A2A protocol URL: https://benchgecko.ai/learn/glossary/a2a-protocol Text reviewed 2026-10-05 TL;DR: A2A (Agent2Agent) is an open protocol, announced by Google in 2025, that lets agents built by different vendors discover each other and exchange tasks. Agents publish a description of what they can do, and other agents send them tasks and receive updates and results over standard web protocols. The project moved to open governance under the Linux Foundation in 2025. ### Compliance #### SOC 2 (SOC 2) URL: https://benchgecko.ai/learn/glossary/soc2 TL;DR: A third-party audit certifying a service provider handles customer data with defined security, availability, and confidentiality controls. SOC 2 is an attestation framework defined by AICPA. Providers are audited against five Trust Services Criteria: Security, Availability, Processing Integrity, Confidentiality, and Privacy. A Type I report covers a point in time; a Type II covers a 6 to 12 month operating period and is what enterprise procurement asks for. Most major AI providers maintain annual SOC 2 Type II reports. Smaller or newer providers without SOC 2 are effectively blocked from regulated verticals (healthcare, finance, government). SOC 2 alone does not cover HIPAA (health) or FedRAMP (US federal), which layer on top. #### FedRAMP URL: https://benchgecko.ai/learn/glossary/fedramp Text reviewed 2026-10-05 TL;DR: A US government authorization required for any cloud or AI service that handles federal agency data. FedRAMP is governed jointly by GSA, DoD, and DHS. Authorization is granted either by a Joint Authorization Board (JAB) or an individual agency. The assessment draws from NIST SP 800-53 controls, with 125+ controls at Moderate and 425+ at High. Cloud providers (AWS GovCloud, Azure Government, Google Cloud for Government) host FedRAMP-authorized AI endpoints. OpenAI offers FedRAMP Moderate through Azure OpenAI Government. Anthropic runs Claude through AWS GovCloud for federal contracts. FedRAMP authorization is a multi-year project · providers without one cannot legally process federal data. #### GDPR (GDPR) URL: https://benchgecko.ai/learn/glossary/gdpr TL;DR: EU regulation governing how personal data of EU residents must be collected, stored, processed, and deleted. GDPR rests on six lawful bases for processing (consent, contract, legal obligation, vital interests, public task, legitimate interests). Personal data includes anything that can identify a person · names, emails, IP addresses, biometric data, and sometimes prompt content. Key rights: access, rectification, erasure ("right to be forgotten"), portability, and objection. Cross-border transfers require Standard Contractual Clauses (SCCs) or adequacy decisions. The EU-US Data Privacy Framework (2023) re-enabled US transfers after Schrems II invalidated Privacy Shield. Frontier AI labs (OpenAI, Anthropic, Google, Mistral) all publish GDPR-aligned DPAs. Mistral defaults to EU data residency as a competitive positioning. #### EU AI Act URL: https://benchgecko.ai/learn/glossary/eu-ai-act Text reviewed 2026-10-05 TL;DR: EU law that classifies AI systems by risk level (unacceptable, high, limited, minimal) and sets obligations for each tier. The AI Act has four risk categories. Unacceptable risk: banned outright. High risk: conformity assessment, risk management system, data governance, documentation, transparency, human oversight, accuracy/robustness/cybersecurity requirements. Limited risk: transparency obligations (disclose AI interaction). Minimal risk: no specific obligations. GPAI providers (OpenAI, Anthropic, Google, Mistral) must provide technical documentation, copyright compliance summaries, and downstream integration guidance. Fines reach €35M or 7% of global revenue · higher than GDPR. #### HIPAA (HIPAA) URL: https://benchgecko.ai/learn/glossary/hipaa Text reviewed 2026-10-05 TL;DR: US federal law protecting patient health data · any AI vendor handling PHI must sign a BAA. HIPAA has two main rules: Privacy Rule (uses and disclosures of PHI) and Security Rule (administrative, physical, technical safeguards for electronic PHI). Violations are tiered by culpability from "did not know" to "willful neglect." The HIPAA Omnibus Rule extended direct liability to business associates. AI providers serving healthcare must offer a BAA · OpenAI offers one for enterprise accounts, AWS for Bedrock healthcare customers, Google Cloud Healthcare API. Logging and model training on PHI are particularly risky; providers typically opt-out of training on healthcare customer data. HIPAA does not preempt stricter state laws (California CMIA). #### Data Residency URL: https://benchgecko.ai/learn/glossary/data-residency TL;DR: The rule that customer data must be stored and processed in a specific country or region, typically the user's jurisdiction. Data residency is distinct from data sovereignty. Residency means "data lives in region X." Sovereignty adds "and is subject only to region X's laws." A US provider with EU-hosted servers satisfies residency but may still face US legal reach (CLOUD Act, FISA). EU customers concerned about sovereignty pick Mistral direct or sovereign clouds (OVH, Scaleway, Deutsche Telekom's DeepL). For AI specifically, residency applies to: prompts at ingestion, generated outputs, logged interactions, and training data if opted in. Each regional endpoint is a separate deployment with its own uptime and model coverage. ### People #### Sam Altman (@sama) URL: https://benchgecko.ai/learn/glossary/sam-altman TL;DR: Sam Altman. Sam Altman is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@sama. Primary topics: AGI, scaling, enterprise. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Dario Amodei (@DarioAmodei) URL: https://benchgecko.ai/learn/glossary/dario-amodei TL;DR: Dario Amodei. Dario Amodei is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@DarioAmodei. Primary topics: safety, scaling, policy. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Yann LeCun (@ylecun) URL: https://benchgecko.ai/learn/glossary/yann-lecun TL;DR: Yann LeCun. Yann LeCun is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@ylecun. Primary topics: open-source, architecture, AGI-skepticism. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Demis Hassabis (@demaborner) URL: https://benchgecko.ai/learn/glossary/demis-hassabis TL;DR: Demis Hassabis. Demis Hassabis is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@demaborner. Primary topics: science, multimodal, nobel. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Jim Fan (@DrJimFan) URL: https://benchgecko.ai/learn/glossary/jim-fan TL;DR: Jim Fan. Jim Fan is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@DrJimFan. Primary topics: embodied-ai, agents, robotics. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Andrej Karpathy (@karpathy) URL: https://benchgecko.ai/learn/glossary/andrej-karpathy TL;DR: Andrej Karpathy. Andrej Karpathy is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@karpathy. Primary topics: education, architecture, agents. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Arthur Mensch (@arthurmensch) URL: https://benchgecko.ai/learn/glossary/arthur-mensch TL;DR: Arthur Mensch. Arthur Mensch is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@arthurmensch. Primary topics: europe, open-source, sovereignty. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Clem Delangue (@ClementDelangue) URL: https://benchgecko.ai/learn/glossary/clem-delangue TL;DR: Clem Delangue. Clem Delangue is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@ClementDelangue. Primary topics: open-source, community, hugging-face. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Swyx (@swyx) URL: https://benchgecko.ai/learn/glossary/swyx TL;DR: Swyx. Swyx is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@swyx. Primary topics: agents, developer-tools, meta-commentary. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Simon Willison (@simonw) URL: https://benchgecko.ai/learn/glossary/simon-willison TL;DR: Simon Willison. Simon Willison is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@simonw. Primary topics: tools, ethics, practical-ai. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Georgi Gerganov (@ggerganov) URL: https://benchgecko.ai/learn/glossary/georgi-gerganov TL;DR: Georgi Gerganov. Georgi Gerganov is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@ggerganov. Primary topics: inference, llama-cpp, quantization. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Harrison Chase (@hwchase17) URL: https://benchgecko.ai/learn/glossary/harrison-chase TL;DR: Harrison Chase. Harrison Chase is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@hwchase17. Primary topics: langchain, agents, rag. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Emad Mostaque (@EMostaque) URL: https://benchgecko.ai/learn/glossary/emad-mostaque TL;DR: Emad Mostaque. Emad Mostaque is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@EMostaque. Primary topics: image-gen, decentralization, controversy. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Aidan Gomez (@aidangomez) URL: https://benchgecko.ai/learn/glossary/aidan-gomez TL;DR: Aidan Gomez. Aidan Gomez is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@aidangomez. Primary topics: enterprise, rag, transformer-origins. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people #### Guillaume Lample (@GusLample) URL: https://benchgecko.ai/learn/glossary/guillaume-lample TL;DR: Guillaume Lample. Guillaume Lample is an AI researcher, founder, or builder tracked by BenchGecko's mindshare layer. On X: @@GusLample. Primary topics: research, mistral, training. Attention metrics (mentions, shares, KOL rank) are tracked daily on /mindshare/people. See live data: https://benchgecko.ai/mindshare/people