Skip to main content
No training required

NVIDIA Releases Kumo Tabular, an Open Foundation Model for Tabular Data

NVIDIA has unveiled Kumo Tabular, an open foundation model that predicts labels for tabular data in a single forward pass—no training, tuning, or feature engineering needed. Available in three sizes and ranked first across four major benchmarks, it applies large language model techniques to enterprise data tables, outpacing gradient-boosted trees and running 17× faster than competitors.
NVIDIA logo and Kumo Tabular text beside a 3D green eye symbol and sample tabular data with numerical and categorical columns.
NVIDIA logo and Kumo Tabular text beside a 3D green eye symbol and sample tabular data with numerical and categorical columns.

NVIDIA has released Kumo Tabular, an open foundation model for tabular classification and regression available on Hugging Face. The model predicts labels for new rows in a single forward pass without training, tuning, or feature engineering. It comes in three sizes ranging from 28M to 215M parameters, runs through NVIDIA's open-source library, and is released under the OpenMDW-1.1 license for commercial use.

Model Size Parameters
Small 28M
Medium ~122M
Large 215M

Kumo Tabular ranks first on four major benchmarks: TabArena, BeyondArena, TALENT, and ScoringBench.

The Problem with Current Approaches

Tabular data forms the backbone of enterprise machine learning. Customer records, transactions, sensor logs, claims, and orders all live in tables, and predicting churn, default, demand, or price from them is among the most common machine learning tasks in industry. For two decades, this work has relied on gradient-boosted trees. But the workflow around those models has barely changed: every new question requires collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that learns each task from scratch.

Large Language Models demonstrated a different approach. Given a few examples in the prompt, a pretrained model solves a task without updating any weights—a technique called in-context learning. This same approach applies to tables: a model pretrained on millions of tables can read a labeled table as context and predict the labels of new rows directly.

Architecture and Design

Kumo Tabular is a Transformer built around table structure, using column, row, and in-context attention as introduced in prior work. To predict a label, the model must: (1) understand what each value means within its column, (2) understand how columns in a row interact, and (3) relate context rows with existing labels to query rows with unknown labels.

Cell Embedding: Groups of cells become tokens. Numerical and categorical values pass through Fourier features—sines and cosines of learned frequencies—with separate weights for each type. Missing values require no imputation and receive special handling. Every token in the context receives a label embedding.

Row Embedding: Rows are embedded by alternating two types of attention. Column attention looks down a single column and learns what a value means in the distribution of its column (whether a value is typical or extreme) via induced self-attention, with computational cost growing linearly with the number of rows. Row attention looks across tokens in a single row to learn feature interactions, using rotary positions to distinguish columns. Four learnable [CLS] tokens join each row as final readout. After row compression, the final stage's cost no longer depends on the number of columns.

In-context Learning: A final Transformer operates on row embeddings. Context rows attend to each other, while query rows attend only to context rows. Each prediction depends only on the context and the row itself, not on which other rows are scored alongside it. Because context never sees queries, its keys and values are computed once and reused for follow-up predictions. Query rows use Test-GQA to shrink the cache. A head turns each query row into class probabilities for classification and 999 quantiles for regression, yielding point predictions and uncertainty estimates.

Length-aware Attention Temperature: Softmax attention spreads as the number of keys grows, potentially becoming unfocused over tens of thousands of rows—a common scenario when inference tables exceed typical training table sizes. Kumo Tabular scales every query by a temperature that grows logarithmically with the number of keys, with a coefficient learned separately for each attention head. This keeps attention sharp as tables grow longer or wider.

Training on Artificial Data

Kumo Tabular is pretrained entirely on artificial tables generated from Structural Causal Models (SCMs). Each training table is sampled in six steps: first by drawing a configuration for table size and task, then by creating a random causal graph evaluated from root to leaf with randomly drawn functions (linear maps, small neural networks, trees, or Gaussian processes). Some nodes become numerical or categorical columns, one becomes the target, and others remain hidden as unmeasured causes. Post-processing correlates column groups, clips outliers, and injects missing values; tables without learnable signal are discarded. The procedural generator produces endless variations, each with new graphs and mechanisms.

Real-world imperfections are built into the generator: values go missing in various patterns, some features are coarsened so duplicate rows may disagree on labels, categorical columns carry many levels, and regression targets can be heavy-tailed. A model trained on millions of such tables learns to handle these issues without cleanup.

Training uses cross-entropy loss for classification and quantile loss for regression, with separate models for each task. Following TabICLv2, training runs in three stages: the first uses tables of 1,024 rows and up to 100 columns; the second varies context from 400 to 10,240 rows; the third extends to 60,000 rows, still with up to 100 columns. In total, Kumo Tabular-Small/Medium/Large saw approximately 35/71/137 million artificial tables. The training recipe and artificial data generators will be released soon.

Model Training Tables (millions)
Small 35
Medium 71
Large 137

Benchmark Results

All three Kumo Tabular sizes were evaluated with default settings against the full TabArena leaderboard, which includes tuned gradient-boosted trees, AutoGluon, and the latest tabular foundation models. Kumo Tabular ranks first overall with an ELO of 1950 while running 17× faster than LimiX-2 under a uniform single RTX 6000 Pro evaluation setup. All three model sizes establish a new state-of-the-art on the accuracy-efficiency Pareto front.

Benchmark Score/Rank
TabArena ELO 1950, first overall
BeyondArena ELO 1418, Improvability 7.78%, first
TALENT (classification accuracy) Average rank 6.67, first overall
TALENT (classification log-loss) Average rank 3.98, first overall
TALENT (regression RMSE) Average rank 4.22, first overall
ScoringBench Large: first, Medium: second on average rank

On BeyondArena, Kumo Tabular achieves an ELO of 1418 with an Improvability score of 7.78%, placing first on the leaderboard. On TALENT, it achieves the top overall ranking across classification accuracy, classification log-loss, and regression RMSE, with average ranks of 6.67, 3.98, and 4.22. On ScoringBench, a benchmark for predictive distributions, Kumo Tabular-Large and Medium rank first and second on average rank.

Limitations and Usage

architecture
architecture — Hugging Face

Kumo Tabular works on numerical and categorical columns only; text, images, or timestamps can be converted to features via built-in pre-processing recipes. A single forward pass covers up to 10 classes, which the library extends to any number of classes with error-correcting output codes. Accuracy may degrade on tables far beyond the training ranges or when query rows come from a different distribution than context rows, so users should validate accuracy and calibration on held-out data before deployment.

The model runs via NVIDIA's newly released GPU-native structured-data-models library, which downloads weights from the Hub on first use and provides preprocessing, ensembling, and many-class handling. A minimal example:

import sdm table = sdm.TableTensor.from_pandas(pd.load_csv(...), device="cuda") na_mask = table["target"].isnan() model = sdm.models.KumoTabular(device="cuda") pred = model( x_context=table[~na_mask].drop_columns("target"), y_context=table[~na_mask, "target"], x_query=table[na_mask].drop_column("target"), )

Kumo Tabular is released under the OpenMDW License Agreement, version 1.1.
Model code and weights are available on GitHub and Hugging Face.

Felipe Santos

“Artificial intelligence can process the world in milliseconds, but only the human heart can give meaning to every second lived” – Mr. Santos