Tabular machine learning · NVIDIA Kumo Tabular explained
Kumo Tabular predicts from labeled table rows without task-specific training
NVIDIA's open model treats labeled rows as context, then returns classifications or numeric predictions in one forward pass. It can remove a training loop from early experiments, but it does not remove the need for held-out tests, calibration checks, or a strong tree-based baseline.
NVIDIA Kumo Tabular is an open foundation model for classification and regression on tables. It uses labeled rows as in-context examples and predicts new rows without task-specific training, tuning, or manual feature engineering. NVIDIA reports leading results across four benchmark suites. Those results are vendor-reported, so practitioners should compare Kumo with tuned tree models and current tabular foundation models on a held-out split before adopting it.
- Tasks
- Classification and regression
- Model sizes
- 28M to 215M parameters
- Pretraining data
- Synthetic tables only
- Weights license
- OpenMDW 1.1
What Kumo Tabular changes
A conventional tabular project trains a separate model for each target, often after feature engineering and hyperparameter search. Kumo Tabular takes a different route. You give it labeled context rows and unlabeled query rows, and it produces class probabilities or regression estimates without updating the model's weights.
That makes it useful as a fast baseline when a team has a clean table and labeled examples but has not built a full training pipeline. It supports classification and regression, includes preprocessing through NVIDIA's structured-data-models library, and comes in three sizes from 28 million to 215 million parameters.
How one forward pass produces a prediction
The model first embeds numerical and categorical cells. Column attention reads values in the context of their column, while row attention learns interactions between fields in one record. A final transformer relates labeled context rows to each query row. Query rows can attend to context rows, but not to other query rows.
For classification, the output is a probability distribution. For regression, the model produces 999 quantiles that support both a point estimate and an uncertainty estimate. NVIDIA also describes key-value caching for the context, which can avoid repeating the same context computation when scoring more rows.
Why synthetic pretraining matters
NVIDIA says Kumo was pretrained on synthetic tables generated from random structural causal models rather than collected customer records. The generator varies causal graphs, functions, missing-value patterns, categorical cardinality, outliers, and regression target shapes. The three model sizes saw about 35 million, 71 million, and 137 million generated tables respectively.
Synthetic pretraining avoids copying a real training table into the weights, but it does not guarantee a match with every business dataset. A table with unusual semantics, leakage, time dependence, or a shifted production population can still fail in ways that synthetic coverage does not predict.
Read the benchmark claims carefully
NVIDIA reports first place on TabArena, BeyondArena, TALENT, and ScoringBench. Its launch article reports a TabArena ELO of 1950 and says Kumo ran 17 times faster than LimiX-2 under a single RTX 6000 Pro setup. These are vendor-reported evaluations, even though the named benchmark projects are public.
A leaderboard average can hide the dataset types that matter to one product. It also does not compare operational concerns such as GPU memory, cold-start time, calibration under distribution shift, or the maintenance cost of a CUDA inference path. Treat the release results as a reason to test, rather than a production guarantee.
Run a held-out comparison before adopting it
- Choose a representative classification or regression task with a frozen train, validation, and held-out test split.
- Compare Kumo with a simple baseline and a tuned gradient-boosted tree using the same rows, target, and leakage controls.
- Measure the task metric, calibration, wall time, peak GPU memory, and failure behavior across missing values and rare categories.
- Repeat the test across several random splits or time windows so one favorable partition does not decide the result.
- Inspect errors for the groups that matter operationally, then test a later or shifted sample before serving predictions.
- Record the exact weights, code revision, preprocessing, ensemble count, hardware, and license terms used in the comparison.
Limits to check first
The native inputs are numerical and categorical columns. NVIDIA says text, images, and timestamps need preprocessing into supported features. One forward pass directly covers up to ten classes, with the library extending larger label sets through error-correcting output codes.
Accuracy may fall on tables outside the model's training ranges or when a query row's distribution differs from its labeled context. The published training stages reached up to 60,000 context rows and 100 columns. The current library requires Python 3.11 or later and PyTorch 2.7 or later, and NVIDIA's examples use CUDA. The code is Apache 2.0, while the model weights use OpenMDW 1.1, so review both before commercial deployment.
Keep learning on Learnetto
Primary sources
NVIDIA launch article on Hugging Face, September 29, 2026