← All AI resources

AI systems engineering · How to Cost Your AI-Powered Filters

A cost model for AI filters—and a better way to reason about query plans

Most AI cost guidance stops at token counts. FS Data Lab works downward into transformer layers and GPU limits, then upward into batching, filter selectivity, short-circuiting, and KV-cache reuse to explain why one AI-SQL plan should beat another.

Published by Learnetto team · Published October 4, 2026

The article derives a speed-of-light latency estimate for AI-powered SQL filters from first principles. For a chosen model and GPU, it counts the work and memory traffic in projections, attention, and MLP layers, then uses filter selectivity and prefix-cache reuse to compare execution orders. Its strongest use is as a planning baseline: it can explain why one plan should beat another before you build it, but real measurements are still required for production capacity and cost decisions.

Source article
Published October 1, 2026 by FS Data Lab
Worked example
Qwen3-4B-FP8 on an NVIDIA H100 SXM
Query
Three AI filters over 5,000 IMDB reviews
Model output
A lower bound on GPU latency, not a price quote

What the article actually does

An AI filter asks a model to return a Boolean judgment for each row or document. That makes a query plan depend on costs that a conventional optimizer does not normally model: transformer weight reads, attention over prompt tokens, GPU arithmetic throughput, HBM bandwidth, and the number of rows that survive each semantic predicate.

FS Data Lab decomposes one forward pass into projections, attention, and the MLP. It applies the roofline model to each component, taking the slower of arithmetic time and memory-transfer time. It then extends that estimate from one filter to a conjunction, where the first filter scans each document and later filters can reuse the document prefix's KV cache.

Key takeaways

  • The model and GPU are part of the query plan. A token count alone cannot tell you latency; the same operation can be limited by arithmetic throughput on one setup and memory bandwidth on another.
  • Batching changes the bottleneck. Larger batches amortize weight reads across more tokens, so projections and MLP layers can move from memory-bound to compute-bound execution.
  • The first AI filter is structurally more expensive because it processes the document prefix. Later filters can ask a shorter question against cached key and value vectors instead of rebuilding the whole prefix.
  • Filter order should balance evaluation cost with rejection power. The best first predicate is not automatically the shortest or the most selective; it is the one that minimizes expected downstream work after accounting for its scan cost.
  • KV-cache reuse is not free. Cache capacity limits how many document prefixes fit in GPU memory, while later attention stages must read saved vectors back from HBM.
  • A speed-of-light estimate is a lower bound, not a forecast. Kernel efficiency, scheduling, framework overhead, host/device transfers or multi-GPU communication when present, and imperfect overlap can make observed latency higher.

How it compares with other approaches

Kipply's Transformer Inference Arithmetic is the closest conceptual companion. It develops the general first-principles intuition behind weight reads, KV-cache capacity, batching, and memory-bound inference. FS Data Lab narrows that machinery to semantic filters and adds database concerns: selectivity, short-circuiting, filter order, and reuse of a shared document prefix.

The FS Data Lab's contrast with empirical profiling describes the opposite route: run representative queries, measure the actual stack, and fit a model. Profiling captures kernel and framework overhead that the roofline estimate omits, so it is usually better for capacity planning on a fixed deployment. Its weakness is portability: change the model, quantization, GPU, runtime, prompt shape, or batching policy and the profile may need to be rebuilt.

LOTUS attacks a broader quality-cost problem. Its semantic operators can use smaller proxy models or embeddings to avoid expensive model work while targeting statistical accuracy guarantees. That changes the physical implementation and may trade a controlled amount of fidelity for speed. FS Data Lab instead estimates the lower-bound cost of executing the chosen model and uses short-circuit ordering that is lossless when predicate evaluations are deterministic and order-independent.

Snowflake's Cortex AISQL is closer to production query optimization. Its engineering explanation shows how it treats model inference as a first-class plan cost, while separate runtime techniques learn model-cascade confidence thresholds and rewrite semantic joins. The FS Data Lab article is more transparent and reproducible as a teaching model; Cortex's approach is more adaptive, but depends on measurements and machinery inside a live engine.

Our take

The valuable idea is not the final latency number. It is the distinction between unavoidable hardware work and avoidable planning work. The gap between measured latency and a credible floor is a diagnostic signal, but it can reflect model simplifications and workload mismatch as well as kernels and serving overhead. The plan model separately shows where reducing rows, reordering predicates, batching more effectively, or reusing prefixes could remove work.

The main caveat is that the worked ordering assumes known selectivities and treats each filter's pass rate as stable regardless of what ran before it. Real predicates can be correlated. A filter that passes half the full dataset may pass nearly everything after another filter, changing the best order. Production use therefore needs sampled conditional selectivities or runtime adaptation, not only global averages.

A practical way to use it

  • Start with one representative query, a fixed model and quantization, one GPU type, and a measured distribution of document and instruction lengths.
  • Calculate the roofline lower bound, then benchmark the same batch shapes on the real serving stack. Keep both numbers: the gap is operational evidence, not an error to hide.
  • Estimate filter selectivity on a labeled sample, including conditional selectivity after likely earlier filters. Do not assume predicates are independent without checking.
  • Compare plausible orders on expected latency and on output equivalence. Short-circuiting is safe for a conjunction only when each predicate's semantics do not depend on execution order.
  • Track quality alongside latency and cost. If you add proxy models, cascades, or approximate retrieval, evaluate the errors they introduce rather than treating them as pure systems optimizations.
  • Recalibrate after changing the model, GPU, quantization, runtime, prompt template, context distribution, or batching policy.

Keep learning on Learnetto

Primary sources

FS Data Lab: How to Cost Your AI-Powered Filters

Kipply: Transformer Inference Arithmetic

Williams et al.: Roofline—An Insightful Visual Performance Model

LOTUS: Semantic Operators and Their Optimization

Snowflake: Optimizing Query Execution in Cortex AISQL

Cortex AISQL research paper