Meta tags:
description= Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.;
Headings (most frequently used words):
index, intelligence, analysis, artificial, per, task, cost, output, aa, speed, vs, model, to, elo, cache, hit, gpt, oss, 120b, high, coding, agent, openness, tokens, time, briefcase, omniscience, image, by, open, weights, cyber, evaluation, leaderboard, provider, arena, preference, 1updated, gdpval, v2, run, pricing, input, and, performance, representation, endpoint, accuracy, independent, of, ai, cybernew, video, speech, capability, indexes, benchmarks, latency, providers, proprietary, release, frontier, language, over, how, read, this, chart, score, text, voice, finance, accounting, evaluations, relevance, v1, components, price, emerging, competition,
Text of the page (most frequently used words):
index (113), intelligence (92), per (69), task (60), #analysis (58), and (55), #artificial (54), cost (52), model (50), the (45), models (40), cache (36), tokens (33), output (33), evaluation (29), new (27), from (25), elo (22), provider (21), reasoning (21), omniscience (21), briefcase (21), not (20), available (20), for (20), input (19), hit (18), better (18), evaluations (17), speed (17), time (17), weighted (17), average (17), publicly (16), sept (16), each (15), higher (15), language (15), price (14), are (14), gdpval (14), terminal (14), bench (14), token (13), usd (13), lcr (13), specific (13), arena (12), see (12), run (12), coding (12), cyber (12), use (11), image (11), gpt (11), more (11), write (11), 691 (11), add (11), answer (11), automationbench (11), humanity (11), last (11), exam (11), gdp (11), pdf (11), performance (10), providers (10), agentic (10), scicode (10), critpt (10), speech (10), methodology (9), api (9), lower (9), how (9), agent (9), video (8), across (8), most (8), endpoint (8), scores (8), weights (8), deepseek (8), openness (8), knowledge (8), first (7), high (7), pricing (7), google (7), accuracy (7), that (7), reference (7), calculated (7), benchmark (7), rate (7), breakdown (7), open (7), sol (7), agents (6), where (6), second (6), prices (6), pareto (6), line (6), attractive (6), quadrant (6), 100 (6), non (6), our (6), meta (6), work (6), hallucination (6), correct (6), includes (6), further (6), details (6), including (6), them (6), incorporates (6), text (6), mistral (6), anthropic (6), openai (6), article (6), published (6), party (5), has (5), with (5), count (5), alibaba (5), kimi (5), xiaomi (5), score (5), tasks (5), spacexai (5), commercial (5), data (4), leaderboard (4), received (4), while (4), streaming (4), oss (4), 120b (4), 500 (4), end (4), latency (4), divided (4), its (4), weight (4), real (4), world (4), leading (4), independent (4), minimax (4), transparency (4), components (4), updated (4), answers (4), means (4), incorrect (4), rubric (4), pass (4), long (4), benchmarks (4), capability (4), oct (4), about (3), lab (3), after (3), which (3), 000 (3), notes (3), offering (3), discount (3), shown (3), here (3), blended (3), other (3), side (3), confidence (3), this (3), relative (3), number (3), all (3), based (3), nvidia (3), flash (3), range (3), measures (3), rewards (3), analytical (3), quality (3), presentation (3), business (3), workflows (3), medical (3), context (3), legal (3), finance (3), votes (3), preference (3), voice (3), top (3), proprietary (3), claude (3), sonnet (3), low (3), your (3), terms (2), 2026 (2), articles (2), recommender (2), optima (2), get (2), figures (2), represent (2), median (2), representation (2), generating (2), chunk (2), been (2), support (2), over (2), cached (2), prompts (2), previously (2), processed (2), typically (2), significant (2), compared (2), regular (2), represented (2), million (2), values (2), storage (2), billed (2), separately (2), vary (2), detail (2), emerging (2), together (2), deepinfra (2), composite (2), bfcl (2), hle (2), 250 (2), against (2), percentage (2), matches (2), point (2), interval (2), dividing (2), comparison (2), used (2), excluding (2), repeats (2), read (2), total (2), segmented (2), type (2), costs (2), one (2), thinking (2), machines (2), 324 (2), availability (2), training (2), different (2), rating (2), max (2), reliability (2), penalizes (2), hallucinations (2), penalty (2), refusing (2), many (2), negative (2), mean (2), than (2), combined (2), metric (2), aggregates (2), head (2), frontier (2), spreadsheets (2), cases (2), may (2), pro (2), under (2), review (2), 2000 (2), multimodal (2), accounting (2), engineering (2), indexes (2), wer (2), editing (2), leaderboards (2), cybergym (2), e2e (2), deepsecbench (2), cwe (2), defense (2), left (2), chart (2), cognition (2), color (2), execution (2), stepfun (2), releases (2), release (2), restricted (2), medium (2), default (2), fallback (2), local (2), gemini (2), argon (2), solar (2), mini (2), large (2), english, privacy, policy, platform, discord, rednote, youtube, linkedin, careers, contact, company, playground, microevals, products, llm, explore, subscribe, email, address, notified, smaller, competitive, competition, scaleway, sambanova, parasail, novita, nebius, base, groq, vertex, turbo, crusoe, coreweave, cerebras, baseten, azure, amazon, measure, much, given, preserves, running, self, hosted, exists, expressed, indicate, lost, quantisation, sampling, defaults, configuration, snapshots, below, within, seconds, decode, minutes, excludes, ttft, overhead, response, log, inverted, stacked, using, compute, required, multiplying, eval, then, post, pre, underlying, contribution, maximum, assesses, basis, their, 283, anchored, 1600, evaluates, economically, valuable, wide, occupations, 566, punishes, bad, guesses, provides, comprehensive, view, produce, factually, reliable, outputs, domains, converted, into, via, synthetic, bounds, clamped, 220, developed, horizon, testing, realistic, require, deliverables, such, presentations, memos, generally, translates, relevant, certain, relevance, mlcr, visual, mmmu, kubernetes, incident, root, cause, itbench, quantitative, documents, analystagent, scientific, research, science, operations, enterpriseops, gym, criterion, harvey, physics, professional, document, saas, writing, faithfulness, instruction, following, user, interaction, private, dataset, tool, measured, independently, economics, healthcare, strategy, ops, capabilities, industries, determined, responses, users, some, due, yet, having, enough, final, transcription, controlled, 167, blind, recruited, human, panel, public, cast, before, january, full, intervals, split, canonical, counts, refused, estimated, safety, blocks, successes, measuring, enterprise, represents, variant, farther, efficient, sit, toward, upper, stronger, results, pay, deepswe, swe, atlas, qna, software, institute, foundation, creators, 501, effort, variants, selected, indicates, whether, labelled, limited, conditions, license, prohibits, country, only, inputs, reaches, xhigh, 0813, 236b, a21b, replaces, just, days, near, astra, agentperf, benchmarking, laptops, workstations, back, three, labs, achieved, korean, upstage, grok, ling, preview, released, making, france, home, intelligent, outside, china, changelog, personalized, recommendations, optimize, priorities, build, own, custom, compare, addressing, urgent, need, step, announcing, understand, landscape, choose, best, case, trends, inference,
Text of the page (random words):
ini 4 argon high new article published 29 sept aa agentperf local benchmarking local ai agents on laptops and workstations new article published 29 sept gpt 6 1 sol replaces gpt 6 sol after just 7 days with near astra intelligence new language model evaluation 29 sept jt 4 1 flash 236b a21b reasoning new language model evaluation 29 sept deepseek v4 pro 0813 non reasoning new language model evaluation 29 sept gpt 6 1 sol medium new language model evaluation 29 sept gpt 6 1 sol low new language model evaluation 29 sept gpt 6 1 sol high new language model evaluation 29 sept gpt 6 1 sol xhigh new language model evaluation 29 sept gpt 6 1 sol max new article published 28 sept claude sonnet 5 5 reaches 2 on the artificial analysis intelligence index new language model evaluation 28 sept claude sonnet 5 5 low default fallback new language model evaluation 28 sept claude sonnet 5 5 medium default fallback see more intelligence coding agent index cyber new image video speech capability indexes benchmarks openness index output tokens cost speed latency providers intelligence intelligence of leading ai models based on our independent evaluations artificial analysis intelligence index artificial analysis intelligence index v4 3 2 incorporates 10 evaluations aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 25 of 691 models add model from specific provider not publicly available artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them open weights proprietary reasoning non reasoning text only multimodal inputs by country artificial analysis intelligence index by open weights proprietary artificial analysis intelligence index v4 3 2 incorporates 10 evaluations aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 25 of 691 models add model from specific provider not publicly available proprietary open weights open weights commercial use restricted artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them open weights indicates whether the model weights are available models are labelled as commercial use restricted if commercial use is limited by conditions and as non commercial if the license prohibits commercial use cost per task time per task output tokens per task cost per intelligence index task weighted average cost usd per artificial analysis intelligence index task segmented by token type lower is better 25 of 691 models not publicly available answer reasoning cache write cache hit input cost per intelligence index task weighted average cost per intelligence index task each evaluation s cost is calculated from input cache hit cache write reasoning and answer token prices divided by task count and weighted by its intelligence index weight intelligence index vs cost per task intelligence index vs time per task intelligence index vs output tokens per task intelligence index vs cost per intelligence index task artificial analysis intelligence index weighted average cost usd per artificial analysis intelligence index task 25 of 691 models most attractive quadrant pareto line stepfun z ai xiaomi minimax openai anthropic mistral meta spacexai nvidia google alibaba deepseek kimi cost per intelligence index task weighted average cost per intelligence index task each evaluation s cost is calculated from input cache hit cache write reasoning and answer token prices divided by task count and weighted by its intelligence index weight artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them intelligence index vs cost per task intelligence index vs time per task intelligence index vs output tokens per task intelligence index vs cost per intelligence index task by model release all reasoning and effort variants of each selected release weighted average cost usd per artificial analysis intelligence index task 10 of 501 releases most attractive quadrant pareto line anthropic openai google meta spacexai xiaomi z ai deepseek other models cost per intelligence index task weighted average cost per intelligence index task each evaluation s cost is calculated from input cache hit cache write reasoning and answer token prices divided by task count and weighted by its intelligence index weight artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them frontier language model intelligence over time artificial analysis intelligence index v4 3 2 incorporates 10 evaluations aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 15 of 58 model creators anthropic openai google meta spacexai xiaomi alibaba z ai stepfun kimi deepseek mistral institute of foundation models minimax thinking machines artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them coding agent index performance cost and execution time for leading coding agents on end to end software engineering tasks index cost execution time artificial analysis coding agent index artificial analysis coding agent index v1 5 incorporates 3 benchmarks deepswe v1 1 terminal bench 4 0 and swe atlas qna higher is better color by model agent 14 of 31 models not publicly available artificial analysis coding agent index vs cost per task artificial analysis coding agent index vs average pay per token api cost per task usd color by model agent 14 of 31 models most attractive quadrant pareto line anthropic google openai anthropic cognition openai cognition spacexai meta z ai kimi alibaba deepseek how to read this chart each point represents a coding agent variant farther left means lower average cost per task while higher on the chart means higher benchmark performance the most efficient agents sit toward the upper left stronger results at lower cost cyber new measuring model capability on enterprise cyber defense cyber index cyber index by benchmark artificial analysis cyber index artificial analysis cyber index v1 incorporates 3 evaluations cwe bench aa deepsecbench aa and cybergym e2e aa successes safety blocks cyber index vs cost per task cwe bench aa vs cost per task deepsecbench aa vs cost per task cybergym e2e aa vs cost per task artificial analysis cyber index score vs cost per task artificial analysis cyber index score vs average cost per task usd lower is better refused tasks are estimated from other models token use most attractive quadrant pareto line z ai xiaomi minimax openai anthropic mistral spacexai nvidia google alibaba meta deepseek kimi evaluation cost per task average cost per task in the evaluation costs are split by input cache hit cache write reasoning and answer token pricing where canonical token counts are available image video top models from our image arena and video arena leaderboards with 95 confidence intervals text to image image editing text to video image to video video editing text to image leaderboard elo scores from blind preference votes by our recruited human panel together with public image arena votes cast before 1 january 2026 see the full leaderboard here 15 of 167 models speech top models from our text to speech arena speech to text and speech to speech evaluations provider voice arena controlled voice arena aa wer index non streaming aa wer streaming index final transcription speech to speech index provider voice arena preference elo arena elo average elo rating of the model higher is better 15 of 96 models arena preference elo relative elo score of the models as determined by responses from users in artificial analysis speech arena some models may not be shown due to not yet having enough votes capability indexes see more measures the performance of models on specific capabilities and industries finance accounting strategy ops legal healthcare medical engineering economics artificial analysis finance accounting index incorporates 7 evaluations aa omniscience gdpval aa v2 1 aa briefcase v1 1 humanity s last exam automationbench aa aa lcr v1 1 gdp pdf higher is better 25 of 25 models not publicly available benchmarks intelligence evaluations intelligence evaluations measured independently by artificial analysis higher is better coding agentic tool use private dataset user interaction finance medical legal intelligence index long context multimodal instruction following faithfulness writing business see more 18 of 27 evaluations 25 of 691 models add model from specific provider not publicly available aa briefcase v1 1 updated agentic knowledge work elo 500 2000 gdpval aa v2 1 updated agentic real world work tasks elo 500 2000 automationbench aa agentic saas workflows terminal bench 4 0 agentic coding terminal use scicode under review coding humanity s last exam reasoning knowledge gdp pdf professional document reasoning all pass critpt under review physics reasoning aa omniscience accuracy knowledge aa omniscience non hallucination rate 1 hallucination rate aa lcr v1 1 long context reasoning harvey lab aa legal agentic work criterion pass rate enterpriseops gym aa agentic business operations terminal bench science 0 1 new agentic scientific research workflows in a terminal aa analystagent quantitative analysis on spreadsheets documents itbench aa kubernetes incident root cause analysis mmmu pro visual reasoning mlcr aa medical long context reasoning intelligence evaluation relevance while model intelligence generally translates across use cases specific evaluations may be more relevant for certain use cases artificial analysis intelligence index artificial analysis intelligence index v4 3 2 includes aa briefcase v1 1 gdpval aa v2 1 automationbench aa terminal bench 4 0 scicode humanity s last exam gdp pdf critpt aa omniscience aa lcr v1 1 see intelligence index methodology for further details including a breakdown of each evaluation and how we run them aa briefcase v1 1 updated aa briefcase is a frontier agentic evaluation for long horizon knowledge work testing agents on realistic business workflows that require deliverables such as spreadsheets presentations and memos aa briefcase elo aa briefcase rubric score analytical quality presentation elo aa briefcase elo vs cost per task aa briefcase elo aa briefcase v1 1 is an agentic knowledge work benchmark developed by artificial analysis aa briefcase elo is a combined metric that aggregates rubric pass rate analytical quality elo and presentation elo higher is better 25 of 220 models add model from specific provider not publicly available aa briefcase elo aa briefcase elo is a combined metric that aggregates analytical quality elo presentation elo and rubric pass rate with rubric performance converted into elo via synthetic head to head matches elo and 95 confidence interval bounds are clamped at 0 aa omniscience aa omniscience is a knowledge and hallucination benchmark that rewards accuracy punishes bad guesses and provides a comprehensive view of which models produce factually reliable outputs across different domains aa omniscience index aa omniscience accuracy aa omniscience hallucination rate aa omniscience index aa omniscience index higher is better measures knowledge reliability and hallucination it rewards correct answers penalizes hallucinations and has no penalty for refusing to answer scores range from 100 to 100 where 0 means as many correct as incorrect answers and negative scores mean more incorrect than correct 25 of 566 models add model from specific provider not publicly available aa omniscience index aa omniscience index higher is better measures knowledge reliability and hallucination it rewards correct answers penalizes hallucinations and has no penalty for refusing to answer scores range from 100 to 100 where 0 means as many correct as incorrect answers and negative scores mean more incorrect than correct gdpval aa v2 1 updated gdpval aa v2 1 evaluates ai models on real world economically valuable tasks across a wide range of occupations gdpval aa v2 1 leaderboard elo rating for performance on real world work tasks anchored to deepseek v4 1 flash max at 1600 higher is better 25 of 283 models add model from specific provider not publicly available openness index artificial analysis openness index assesses how open models are on the basis of their availability and transparency across different components openness index components openness index artificial analysis openness index components openness index underlying score contribution by components up to a maximum of 18 higher is more open 10 of 324 models add model from specific provider transparency pre training data transparency post training data transparency methodology model availability artificial analysis openness index vs artificial analysis intelligence index 10 of 324 models add model from specific provider most attractive quadrant pareto line xiaomi z ai kimi deepseek alibaba minimax thinking machines nvidia meta output tokens output tokens of leading ai models based on our independent evaluations output tokens per task intelligence index vs output tokens per task output tokens intelligence index vs output tokens output tokens per intelligence index task weighted average number of output tokens used to run one task in the artificial analysis intelligence index 25 of 691 models not publicly available answer reasoning output tokens per intelligence index task the number of tokens required per intelligence index task this is calculated by multiplying the output tokens per eval by the relative weights of each benchmark in the intelli...
|