NVIDIA Kumo Tabular release achieves state-of-the-art accuracy‑efficiency on tabular benchmarks
TL;DR
NVIDIA Kumo Tabular is an open‑source foundation model for tabular data that predicts labels in a single forward pass without any fine‑tuning, and it ranks first on the TabArena, BeyondArena, TALENT, and ScoringBench benchmarks, establishing a new accuracy‑efficiency frontier.
The Shift to Tabular Foundation Models
Tabular data underpins most enterprise machine‑learning workloads, yet traditional pipelines still rely on gradient‑boosted trees that require bespoke feature engineering, hyper‑parameter search, and full retraining for each new task. Inspired by in‑context learning in large language models, Kumo Tabular treats a labeled table as a prompt and directly predicts labels for new rows, eliminating the need for any weight updates.
"Given a table with labeled rows and the rows you want predictions for, Kumo Tabular returns class probabilities or numeric predictions in a single forward pass." – NVIDIA Kumo Tabular announcement
How Kumo Tabular Works
Kumo Tabular adapts the Transformer architecture to the intrinsic structure of tables using three key mechanisms:
- Cell Embedding – Numerical and categorical values are transformed with Fourier features (sine/cosine of learned frequencies). Missing values receive a dedicated token, and each context cell is paired with a label embedding.
- Row Embedding – Two alternating attention layers capture column‑wise distributions (column attention) and intra‑row feature interactions (row attention). Four learnable
[CLS]tokens per row serve as the final row representation. - In‑Context Learning – A top‑level Transformer processes row embeddings: context rows attend to each other, while query rows attend only to the context. Test‑GQA caching reduces the computation needed for each query. The output head produces class probabilities for classification and 999 quantiles for regression, yielding point predictions and uncertainty estimates.
Length‑Aware Attention Temperature
To keep attention sharp as table size grows, Kumo Tabular scales the softmax temperature with the logarithm of the number of keys, using a head‑specific learned coefficient. This preserves discriminative power for tables that are orders of magnitude larger than those seen during pre‑training.
Pre‑Training on Artificial Tables
Kumo Tabular is trained exclusively on synthetic tables generated from structural causal models (SCMs):
- Random causal graphs define hidden variables and target relationships.
- Nodes are instantiated as numerical or categorical columns via diverse functions (linear maps, small neural nets, trees, Gaussian processes).
- Post‑processing adds realistic imperfections—missingness patterns, coarsened features, heavy‑tailed targets, and high‑cardinality categories.
- Tables lacking a learnable signal are filtered out using a quick tree‑ensemble check.
Training proceeds in three stages, scaling context size from 1,024 rows (stage 1) to up to 60,000 rows (stage 3) while keeping column count ≤ 100. The three model sizes—Small (28 M), Medium (≈ 100 M), Large (215 M) parameters—saw roughly 35 M, 71 M, and 137 M synthetic tables respectively.
Benchmark Performance
Evaluated under a uniform RTX 6000 Pro setup, all three Kumo Tabular variants outperform existing tabular foundation models and tuned gradient‑boosted trees on four major leaderboards:
- TabArena – Kumo Tabular achieves the highest overall ELO (1950) and runs 17 × faster than the previous state‑of‑the‑art LimiX‑2.
- BeyondArena – ELO of 1418 and an Improvability score of 7.78 %, placing first.
- TALENT – Top average ranks for classification accuracy (6.67), classification log‑loss (3.98), and regression RMSE (4.22).
- ScoringBench – Kumo Tabular‑Large and Medium rank first and second on average predictive‑distribution rank.
The results place Kumo Tabular on the new accuracy‑efficiency Pareto frontier, as illustrated in the accompanying benchmark plots.
Limitations
- Supports only numerical and categorical columns; other modalities (text, images, timestamps) must be pre‑processed into features.
- Single‑pass inference handles up to 10 classes directly; larger label spaces require error‑correcting output codes provided by the library.
- Accuracy may degrade on tables that exceed the training range or when query rows come from a distribution different from the context rows. Validation on held‑out data is essential before deployment.
Quick Demo
The NVIDIA structured-data-models library provides a GPU‑native interface. The following Python snippet converts a pandas.DataFrame to a TableTensor, supplies context rows with labels, and obtains predictions for missing‑label rows:
import sdm # structured-data-models
import pandas as pd
# Load a CSV into a DataFrame and tensorize it on the GPU
table = sdm.TableTensor.from_pandas(pd.read_csv('data.csv'), device='cuda')
na_mask = table['target'].isnan()
model = sdm.models.KumoTabular(device='cuda')
pred = model(
x_context=table[~na_mask].drop_columns('target'),
y_context=table[~na_mask, 'target'],
x_query=table[na_mask].drop_column('target'),
)
The library automatically downloads the pretrained weights from the Hugging Face Hub on first use and handles preprocessing, ensembling, and multi‑class extensions.
Getting Started
Kumo Tabular is released under the OpenMDW‑1.1 license, permitting commercial use. Developers should follow NVIDIA’s Trustworthy AI policies, validate model performance on domain‑specific data, and report any quality, risk, or security concerns via the GitHub issue tracker.
- Model code: https://github.com/NVIDIA/structured-data-models
- Model weights: https://huggingface.co/nvidia/Kumo-Tabular
Acknowledgements
The authors thank David Holzmüller for significant ideas and ablations, and Vignesh Kothapalli for his internship contributions.