benchmark measured on an RTX 4070 Ti


The declare has been going round for months and it’s concrete sufficient to be measurable: a tabular basis mannequin predicts on a desk with out ever having skilled on it and nonetheless beats tuned boosting.

If that’s true, half a decade of observe adjustments form. Searching hyperparameters stops being a compulsory step and turns into a luxurious that generally doesn’t repay.

So I put it to the check alone card, with fourteen datasets, 4 contenders and the identical stopwatch for everybody.

In 33 seconds with narration: two bots compete over a desk. The one-eyed one seems and solutions; the geared one tries twenty-five combos earlier than replying. Every quantity is a measured one. Muted by default: flip it on within the controls.Watch it in the reel viewer →

What precisely these fashions do

A tabular basis mannequin is pretrained on thousands and thousands of artificial tables generated on function. When a brand new desk arrives, it doesn’t modify a single weight: it receives the coaching rows as context and produces the predictions in a single ahead cross.

It is identical in-context studying concept we already know from language fashions, moved from phrases to columns. That is why the verb “prepare” sits oddly: the code nonetheless calls it match, however inside there is no such thing as a gradient descent, there’s a copy of knowledge to the cardboard.

That additionally explains why the associated fee reveals up the place you don’t anticipate it. Fitting is almost free and prediction is what pays, precisely the reverse of a tree.

Put that method it sounds summary, so right here is the complete journey of 1 row: a request arrives with an empty cell, the mannequin weighs it in opposition to every part that already occurred, resolves it in a single cross and returns a future. Pick any of the six instances I measured to see it with their actual knowledge:

The request

The context

A single cross

The prediction

A credit score utility arrives

2 100 functions already settled · 10 columns
clf_num/credit score.csv

TabICL · nothing tuned · 0.8 s

Will they repay?

repaysdefaults

AUC 0.8528
credit score

A contact enters the record

2 100 calls already made · 7 columns
clf_num/bank-marketing.csv

TabICL · nothing tuned · 0.8 s

Is it value calling them?

indicators updoesn’t

AUC 0.8699
bank-marketing

A affected person is discharged

2 100 earlier discharges · 7 columns
clf_num/Diabetes130US.csv

TabICL · nothing tuned · 0.8 s

Will they be readmitted?

returnsdoesn’t

AUC 0.6484
Diabetes130US

A case is assessed

2 100 instances already closed · 11 columns
clf_cat/compas-two-years.csv

TabICL · nothing tuned · 0.6 s

Will they reoffend?

reoffendsdoesn’t

AUC 0.7329
compas-two-years

A market interval closes

2 100 earlier intervals · 7 columns
clf_num/electrical energy.csv

TabICL · nothing tuned · 0.8 s

Does the worth go up or down?

updown

AUC 0.8873
electrical energy

A candidate molecule arrives

2 100 molecules already assayed · 419 columns
clf_num/Bioresponse.csv

TabICL · nothing tuned · 6.0 s

Does it set off a organic response?

livelyinert

AUC 0.8667
Bioresponse

The request

The context

A single cross

The prediction

A credit score utility arrives

2 100 functions already settled · 10 columns
clf_num/credit score.csv

TabICL · nothing tuned · 0.8 s

Will they repay?

repaysdefaults

AUC 0.8528
credit score

A contact enters the record

2 100 calls already made · 7 columns
clf_num/bank-marketing.csv

TabICL · nothing tuned · 0.8 s

Is it value calling them?

indicators updoesn’t

AUC 0.8699
bank-marketing

A affected person is discharged

2 100 earlier discharges · 7 columns
clf_num/Diabetes130US.csv

TabICL · nothing tuned · 0.8 s

Will they be readmitted?

returnsdoesn’t

AUC 0.6484
Diabetes130US

A case is assessed

2 100 instances already closed · 11 columns
clf_cat/compas-two-years.csv

TabICL · nothing tuned · 0.6 s

Will they reoffend?

reoffendsdoesn’t

AUC 0.7329
compas-two-years

A market interval closes

2 100 earlier intervals · 7 columns
clf_num/electrical energy.csv

TabICL · nothing tuned · 0.8 s

Does the worth go up or down?

updown

AUC 0.8873
electrical energy

A candidate molecule arrives

2 100 molecules already assayed · 419 columns
clf_num/Bioresponse.csv

TabICL · nothing tuned · 6.0 s

Does it set off a organic response?

livelyinert

AUC 0.8667
Bioresponse

Default threat

the inspiration mannequin wins

Banking’s most repeated case: deciding who will get lent to.

dataset dimension TabICL TabPFN tuned XGB
credit score 3000 × 10 0.7667 0.7578 0.7533
heloc 3000 × 22 0.7222 0.7300 0.7078
default-of-credit 3000 × 20 0.6956 0.6967 0.6944

All three datasets go to the inspiration mannequin. On heloc TabPFN takes 0.022 and TabICL 0.014: in credit score threat that’s not ornament.

Context: clf_num/credit.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 22a759296600c39a56884dafd84eb346be018c537ff096adc8c2ae7e0520f2a9

Sales marketing campaign

the inspiration mannequin wins

Who to name first when there are a thousand contacts and time for 100.

dataset dimension TabICL TabPFN tuned XGB
bank-marketing 3000 × 7 0.7944 0.7967 0.7833

The solely case the place TabPFN finally ends up forward of TabICL. Both beat tuned boosting.

Context: clf_num/bank-marketing.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 3433fecd416ce949692442882669881746e8a6b24b904a5f808cdcb196443e7e

Clinical readmission

the inspiration mannequin wins

An actual hospital document: which affected person will get readmitted.

dataset dimension TabICL TabPFN tuned XGB
Diabetes130US 3000 × 7 0.5911 0.5933 0.5900

On accuracy they almost tie, however the space below the curve opens up sharply: 0.6484 in opposition to 0.6264. The threat ordering, which is what a triage makes use of, improves significantly greater than accuracy suggests.

Context: clf_num/Diabetes130US.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 9be384ad7edbb9a98b509adb9cc08578e63bd87545b05d97d5147aa15b71385a

Recidivism

the inspiration mannequin wins

The dataset that opened the controversy on algorithmic bias within the courts.

dataset dimension TabICL TabPFN tuned XGB
compas-two-years 3000 × 11 0.6733 0.6767 0.6700

The basis mannequin wins, and it’s value saying that a greater mannequin right here doesn’t make the use professional: this dataset’s argument was by no means about accuracy.

Context: clf_cat/compas-two-years.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 7e0bcef09ae633ad81e26b0a8f8f96dbc4ac3c29b9da0760d38ef2a75adde8d8

Electricity demand

a tie

Consumption and value sequence from an electrical energy market.

dataset dimension TabICL TabPFN tuned XGB
electrical energy 3000 × 7 0.8178 0.8111 0.8178

The tie. TabICL matches tuned boosting to the fourth decimal and TabPFN falls beneath. It is likely one of the two instances the place the benefit doesn’t present up.

Context: clf_num/electricity.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 d6005007c4b1f7ba88cf91cb1f231969cce3e49c98d445338c0ab0342ead5d7e

Molecular screening

relies on the mannequin

419 columns of chemical descriptors: the broad desk.

dataset dimension TabICL TabPFN tuned XGB
Bioresponse 3000 × 419 0.7922 0.7567 0.7744

This is the place TabPFN breaks: 0.7567 is worse than even untuned XGBoost, and it prices 18.1 seconds. TabICL holds up and comes first. The desk’s width, not its size, is what squeezes.

Context: clf_num/Bioresponse.csv from the inria-soda/tabular-benchmark suite, trimmed to three,000 rows with seed 0 and break up 70/30 with stratification.
sha256 bd3d277821eb949df41219363f0386549019249776fccf7a9f08d3bec10b8727

Pick a case above to comply with one row’s journey. The context is the two 100 coaching rows, the time is the one TabICL measured, and the orb’s arc attracts that dataset’s space below the curve, from probability to an ideal rating.

What that is for, concretely

The six instances within the diagram usually are not brochure examples: they’re the datasets I measured with. But it’s value widening the map, as a result of “tabular basis mannequin” appears like a laboratory and the issue it solves is likely one of the most typical there may be.

A desk is a spreadsheet: rows which might be instances and columns which might be attributes. And the duty is all the time the identical, predict one column from the others:

  • A buyer desk with tenure, plan, utilization and complaints, to estimate which of them are going to depart subsequent month.
  • A transaction log with quantity, service provider, hour and nation, to flag which of them are fraud.
  • A historical past of credit score functions with earnings, debt and cost behaviour, to estimate who shouldn’t be going to repay. Three of the datasets I measured are precisely that: credit score, heloc and default-of-credit.
  • Sensor readings from a machine, to anticipate when it’ll break.
  • Patient information with signs and lab outcomes, to prioritize who will get seen first. Diabetes130US, one other of the datasets, is an actual hospital document.

That is half the information work performed in any firm. It can be the bottom the place deep studying had been dropping for years: for tables, a great gradient-boosted tree ensemble was nonetheless the appropriate reply, and the analysis that assembled these datasets was titled precisely that method, asking why bushes nonetheless received.

What adjustments if the promise holds

Today, fixing any of these instances has a ritual: put together the information, select a mannequin, search hyperparameters, validate, repeat. The search is the boring half and the one which eats machine and human hours. In this benchmark, XGBoost’s search took as much as 50 seconds per dataset; on an actual downside with extra rows and extra combos, it’s minutes or hours.

A tabular basis mannequin proposes skipping that ritual totally. You hand it the desk, it solutions in a second, and that’s that. No tree depth to decide on, no studying fee, no cross-validation to resolve amongst twenty-five candidates.

Put in concrete phrases: as an alternative of spending the afternoon tuning a churn mannequin, you’ve got a solution within the time it takes to make espresso, and solely then do you resolve whether or not something is value refining. For exploring a desk that simply arrived, or for having an sincere baseline earlier than investing time, it’s onerous to beat.

The query, then, shouldn’t be whether or not the concept is engaging. It is whether or not the end result holds up when measured.

The benchmark’s guidelines

A badly constructed benchmark says no matter you need. These are the principles I imposed on myself earlier than seeing a single quantity:

The knowledge is third-party and from the sphere the place this argument is fought. I used Grinsztajn’s tabular benchmark, the suite gathered for the work asking why bushes nonetheless beat deep studying on tables. Choosing the datasets your self is the simplest solution to manufacture the end result you want.

XGBoost goes in tuned, not for adornment. The declare says “tuned boosting”, so evaluating in opposition to manufacturing facility parameters can be a straw man. I ran two variations: one with mounted, affordable parameters and one other with a random search over twenty-five combos and three-fold cross-validation.

Same break up, similar seed, similar columns for everybody. Categoricals are integer-encoded. It shouldn’t be the very best remedy, however it’s similar for all 4, which is what makes the quantity comparable.

Five seeds per dataset, and the median is reported. A single break up can’t inform sign from luck.

Prediction time consists of two inference passes, the category one and the chance one, as a result of the benchmark wants each to compute accuracy and space below the curve. That makes issues dearer exactly for the inspiration fashions, which is the place their value lives, so the seconds I publish for them are inflated in no person’s favour.

Two cores are left free. The machine has different work on it, and a benchmark that eats the entire CPU measures the struggle with the scheduler, not the mannequin.

Everything ran on a 16 GB RTX 4070 Ti SUPER, with fourteen cores for XGBoost.

Three stumbles earlier than the primary quantity

Publishing solely the ultimate desk can be mendacity by omission. This is what it value to get there.

PyTorch turned the GPU off with out saying so

I arrange the surroundings, ran the benchmark, and the primary line mentioned system: cpu. The card was free and visual. The motive:

2.13.0+cu130   driver 570.144   torch.cuda.is_available() → False

PyTorch had resolved to a construct for CUDA 13.0 whereas the machine’s driver exposes 12.8. Instead of failing, it silently turns the GPU off and carries on with the CPU. The warning solely reveals up when you examine torch.cuda.is_available() by hand.

It is mounted by pinning the model to the appropriate index:

pip set up --no-cache-dir 
  --index-url https://download.pytorch.org/whl/cu128 torch==2.9.1+cu128

This is the second time this 12 months the identical entice has value me a complete run. If a GPU benchmark offers suspiciously gradual numbers, that’s the first place to look.

OpenML’s API returned 504 on each endpoint

The unique plan was to take the datasets from OpenML, which is the canonical supply for these comparisons. It failed fully in the course of the run:

https://api.openml.org/api/v1/json/data/31          → 504 (16.9 s)
https://www.openml.org/api/v1/json/data/31          → 504 (15.7 s)
https://api.openml.org/api/v1/json/data/features/31 → 504 (16.5 s)

Four retries with rising backoff made no distinction: it was not a fee downside, it was the gateway being down. I switched the supply to the identical benchmark’s CSVs hosted on HuggingFace, which resolved in 250 milliseconds. It is an uncomfortable reminder of how a lot reproducible analysis relies on a single service.

The most-cited mannequin not downloads with out an account

This is the discovering that pursuits me most, as a result of it isn’t technical.

TabPFN is the identify that seems in almost all protection of this subject. I put in the present model, 8.3.0, and on the primary match:

TabPFNLicenseError: TabPFN requires a one-time license acceptance
to obtain mannequin weights for native inference, however no interactive
terminal is on the market.

To fetch the weights it’s important to open a browser, register, settle for the license in a tab on the seller’s website and export an account token. The best-known tabular basis mannequin stopped being one thing you put in and run.

There are two methods out, and I attempted each:

  • TabPFN’s 2.x sequence nonetheless downloads the weights with out registering. I put in 2.2.1 in a separate surroundings and it labored first strive. That is the one within the tables beneath.
  • TabICL, from Inria’s Soda group, has a three-clause BSD license and downloads with no barrier in any respect. It additionally turned out to be the higher of the 2.
The benchmark’s two contenders: on the left the one which simply seems on the desk and solutions, on the appropriate the one which spends half a minute making an attempt combos.

The numbers

Fourteen datasets, all trimmed to three,000 rows so the comparability is homogeneous. Accuracy because the median of 5 seeds. The seconds column sums becoming and prediction, which is every one’s sincere value:

dataset rows × cols TabICL TabPFN 2.2.1 XGBoost Tuned XGB s TabICL s tuned XGB
bank-marketing 3000 × 7 0.7944 0.7967 0.7711 0.7833 0.8 1.8
credit score 3000 × 10 0.7667 0.7578 0.7478 0.7533 0.8 2.2
heloc 3000 × 22 0.7222 0.7300 0.7022 0.7078 0.8 1.6
pol 3000 × 26 0.9844 0.9833 0.9778 0.9756 0.8 1.0
eye_movements 3000 × 20 0.6100 0.6111 0.5944 0.5833 0.8 3.0
california 3000 × 8 0.8967 0.8978 0.8833 0.8767 0.6 1.6
house_16H 3000 × 16 0.8811 0.8744 0.8722 0.8711 1.1 3.0
MagicTelescope 3000 × 10 0.8778 0.8622 0.8489 0.8456 1.3 2.4
electrical energy 3000 × 7 0.8178 0.8111 0.8156 0.8178 0.8 2.4
Diabetes130US 3000 × 7 0.5911 0.5933 0.5644 0.5900 0.8 1.4
default-of-credit 3000 × 20 0.6956 0.6967 0.6878 0.6944 0.9 4.5
compas-two-years 3000 × 11 0.6733 0.6767 0.6444 0.6700 0.6 0.7
Bioresponse 3000 × 419 0.7922 0.7567 0.7844 0.7744 6.0 27.2
albert 3000 × 31 0.6500 0.6611 0.6511 0.6611 1.8 2.6

Against tuned XGBoost, by accuracy:

  • TabICL: twelve wins, one tie (electrical energy) and one loss (albert). Mean distinction +0.0106.
  • TabPFN 2.2.1: eleven wins, one tie (albert) and two losses (electrical energy and Bioresponse). Mean distinction +0.0075.

Accuracy, nonetheless, is a rough metric: it relies on the brink and punishes in a different way relying on how the lessons are balanced. Area below the curve is extra informative, and there the end result will get sharper:

mannequin AUC wins imply distinction
TabICL 14 of 14 +0.0114
TabPFN 2.2.1 13 of 14 +0.0089

TabICL beats tuned XGBoost on each dataset with out exception. Even on albert, the place it loses on accuracy, it has a greater AUC (0.7101 in opposition to 0.7094). That element is precisely why each metrics are value taking a look at: accuracy mentioned “loss” the place the chance ordering mentioned “win by a hair”.

That mentioned, the passion wants calibrating. Winning fourteen of fourteen is a powerful sign of consistency, not of crushing superiority: the imply distinction is one hundredth. Nobody goes to note that on a dashboard. What does get seen is the opposite factor.

The value is the reverse of what you anticipate

Look once more on the desk’s final two columns. TabICL resolves a dataset in below a second with out tuning something. Tuned XGBoost takes between 0.7 and 27.2 seconds looking hyperparameters, and nonetheless comes out beneath.

That is the actual argument, and it isn’t accuracy. It is that the costly a part of the work, the one which consumes human and machine time, merely disappears.

The shock: the benefit doesn’t break with dimension

This is the place I anticipated to dismantle the promise. The normal objection to those fashions is that they solely work on toy tables, as a result of the coaching rows have to suit contained in the transformer’s context.

I took jannis, which has 57,580 rows, and trimmed it to rising sizes. Three seeds per dimension:

rows TabICL XGBoost Tuned XGB s TabICL s tuned XGB
500 0.7533 0.7733 0.7533 0.7 3.2
1,000 0.7733 0.7467 0.7533 0.9 4.9
2,000 0.7800 0.7583 0.7667 1.1 9.4
4,000 0.7817 0.7692 0.7725 1.8 21.1
8,000 0.7958 0.7658 0.7583 2.6 19.7
16,000 0.8090 0.7785 0.7852 5.3 34.6
32,000 0.8230 0.7894 0.7929 11.7 50.3

The benefit doesn’t break, and on the giant finish it consolidates: +0.030 at 32,000 rows. And the break up of occasions opens within the path reverse to instinct: 11.7 seconds in opposition to 50.3.

It is value saying exactly what grows and what doesn’t, as a result of the distinction issues. What rises steadily is TabICL’s absolute accuracy: 0.7533, 0.7733, 0.7800, 0.7817, 0.7958, 0.8090, 0.8230, monotone throughout all seven measurements. The benefit over XGBoost, in contrast, is irregular: 0.000, +0.020, +0.013, +0.009, +0.038, +0.024, +0.030. It rises and falls.

And at 4,000 rows that +0.009 benefit is smaller than the unfold throughout the three seeds (TabICL ranges from 0.768 to 0.799; tuned XGBoost from 0.758 to 0.790), so at that time it’s indistinguishable from noise and I don’t depend it as a win.

Where it’s strong is on the high: at 32,000 rows TabICL’s worst seed (0.8216) sits above XGBoost’s greatest (0.7950). There is not any doable overlap there.

The solely place XGBoost wins cleanly is the small finish, at 500 rows, the place the variability between seeds is so excessive that I’d not wager something on that distinction.

If you have been anticipating, as I used to be, the curve to flip sooner or later, it doesn’t accomplish that right here. You must go significantly greater to search out it.

Illustration: a robot straining to push an enormous wall of glowing data columns.
Bioresponse has 419 columns. The desk’s size doesn’t cease them; its width does.

Where they do break

Bioresponse is the telltale dataset: 419 columns.

TabPFN 2.2.1 falls to 0.7567, the worst of the 4 contenders, beneath even untuned XGBoost. And it prices it 18.1 seconds, twenty-five occasions greater than on a standard dataset. TabICL holds up significantly better (0.7922, the perfect of the 4) however pays too: 6.0 seconds in opposition to the standard 0.8.

The studying is that the desk’s width, not its size, is the axis that genuinely squeezes. That is smart: the variety of columns enters the transformer’s consideration value, and the artificial pretraining covers tables of tens of columns effectively, not tons of.

What this measurement doesn’t say

I’d moderately record the bounds than faux they don’t exist:

  • Binary classification solely. I didn’t check regression or multiclass.
  • A 3,000-row ceiling in the principle desk. The scale sweep reaches 32,000, however on a single dataset.
  • The hyperparameter search was twenty-five combos. More aggressive tuning would shut a part of the hole; how a lot, I have no idea, as a result of I didn’t measure it.
  • XGBoost’s search optimized accuracy, and afterwards I additionally examine by space below the curve. Which means the “fourteen of fourteen on AUC” is in opposition to a boosting mannequin that was not tuned for that metric. Tuning it for AUC would in all probability enhance it there; I didn’t measure that.
  • Categoricals have been integer-encoded for everybody. Better remedy would favour XGBoost greater than the others.
  • The encoding and null-filling have been computed over the entire desk, earlier than splitting. It is identical leak for all 4 fashions, so it doesn’t change who wins, nevertheless it inflates everybody barely and shouldn’t be performed that method.
  • Several accuracy wins are smaller than the variation between seeds. Diabetes130US is received by 0.0011 when the per-seed distinction ranges from −0.011 to +0.049: there the signal relies on which seed comes up. The area-under-the-curve depend does maintain seed by seed; the accuracy one, on three or 4 datasets, doesn’t.
  • The wide-table discovering rests on a single dataset. Bioresponse is the one one with tons of of columns, so “width is what squeezes” is a speculation with one remark, not a rule.
  • The present model of TabPFN, 8.3.0, went unmeasured due to the license barrier. TabPFN’s numbers are from 2.2.1, two sequence behind.
  • The CPU was shared with different work on the machine. That impacts XGBoost’s occasions greater than the GPU’s, so if something it performs in opposition to the inspiration fashions within the timing comparability.

When I’d use it and when not

I’d use it on any desk between one thousand and thirty thousand rows with fewer than 100 columns, the place an individual’s time is value greater than the final hundredth. A second of compute, zero tuning, and a end result that in my measurement was constantly higher. For an preliminary exploration it’s onerous to justify not doing it.

I’d not use it on very broad tables, the place it degrades and will get costly. Nor in manufacturing and not using a GPU, nor the place you want a small artifact that may be inspected and deployed in a light-weight container: a skilled tree weighs kilobytes and runs anyplace, whereas right here it’s important to load a transformer.

And if the factors embody having the ability to audit the place the weights got here from, as we speak the reply is TabICL. Not for efficiency, although it was additionally the higher of the 2, however as a result of it’s the one you possibly can nonetheless obtain and run with out asking anybody’s permission.


Measured on 16 August 2026 on a 16 GB RTX 4070 Ti SUPER, with torch 2.9.1+cu128, XGBoost 3.4.1, TabICL 2.1.1 and TabPFN 2.2.1. Data from the inria-soda/tabular-benchmark tabular suite. The important desk took 585 seconds and the size sweep 514.

Sources



Source link