GPT-6.1 Sol replaces GPT-6 Sol after simply 7 days, with near-Astra intelligence


See model page

GPT-6.1 Sol replaces GPT-6 Sol after simply 7 days. It scores 1 level under GPT-6 Astra within the Intelligence Index at lower than one quarter of the Cost per Task.

Pricing matches GPT-6 Sol at $2/$10 per million enter/output tokens, besides that the cache learn low cost rises from 90% to 95%. GPT-6.1 Sol’s general blended value for agentic workloads is subsequently barely decrease than GPT-6 Sol. This represents a further value minimize, following GPT-6 Sol’s authentic 50% low cost from GPT-5.6 Sol.

Key takeaways:

➤ Achieves near-Astra Intelligence: GPT-6.1 Sol features 4 factors within the Intelligence Index vs GPT-6 Sol, and 5 factors vs GPT-5.6 Sol – touchdown 1 level under GPT-6 Astra. It makes important features in agentic information work, enhancing 4 factors and 5 factors in AA-Briefcase v1.1 and GDPval-AA v2.1 respectively. Other notable features embody a 12 level leap in Terminal-Bench 4.0, a 5 level leap in Humanity’s Last Exam, a 6 level leap in GDP.pdf, and an 8 level leap in AA-Omniscience Accuracy coupled with hallucination price falling from 60% to 54%.

➤ Pushes price effectivity frontier: At max effort, GPT-6.1 Sol prices lower than 1 / 4 of GPT-6 Astra per Intelligence Index job ($0.72 vs $3.26). It additionally prices 31% much less per job than GPT-6 Sol ($1.05) and 64% lower than GPT-5.6 Sol ($1.99). All effort ranges of GPT-6.1 Sol push out the associated fee effectivity Pareto frontier: for a given stage of intelligence, there isn’t any cheaper mannequin.

➤ Pushes token effectivity frontier, however makes use of barely extra output tokens than GPT-6 Sol: GPT-6.1 Sol makes use of ~10-30% extra output tokens than GPT-6 Sol throughout effort ranges. However, as a result of enhance in Intelligence Index rating, its low and medium effort ranges are Pareto optimum for token effectivity.

➤ Gains in Coding Agent Index: GPT-6.1 Sol features 3 factors on GPT-6 Sol at max effort within the Artificial Analysis Coding Agent Index, and sits 2 factors under GPT-6 Astra.

Coding Agent Index

GPT-6.1 Sol dominates the lower-price vary of the Pareto frontier for Artificial Analysis Coding Agent Index vs Cost per Task. GPT-6.1 Sol (xhigh) scores 1 level above GPT-6 Astra for lower than 15% of the Cost per Task. This represents a 6 level achieve from GPT-6 Sol (max). We noticed the xhigh effort setting to outperform the max effort setting by 3 factors.

Token effectivity

GPT-6.1 Sol makes use of 10-30% extra output tokens than GPT-6 Sol throughout effort settings within the Intelligence Index. However, as a result of will increase in intelligence, its low and medium effort settings are Pareto optimum for token effectivity.

AA-Omniscience

At max effort, GPT-6.1 Sol jumps 8 factors in AA-Omniscience Accuracy coupled with a 6 level discount in hallucination price.

AA-Briefcase

GPT-6.1 Sol improves by ~80 Elo in AA-Briefcase. This is pushed by will increase in its rubric rating and Analytical Quality Elo, whereas Presentation Elo falls barely.

Results by analysis

Breakdown of the person evaluations within the Artificial Analysis Intelligence Index v4.3.2.

Compare GPT-6.1 Sol with different main fashions at: artificialanalysis.ai/models/releases/gpt-6-1-sol



Source link