All AI videos

AI video

Small Models, Big Results: Training a Finance Agent for Under $500 — Charles Dickens, Snorkel AI

Snorkel AI researcher Charles Dickens presents a reinforcement-learning recipe for a 4B financial agent and the published benchmark results used to compare it with larger models.

AI Engineer · 2026 featured video

Watch on YouTube

What this video shows

Dickens describes a financial question-answering agent with table-listing, schema, SQL, and calculator tools. The team starts from Qwen3-4B-Instruct and trains with reinforcement learning on verified questions drawn from SEC 10-K filings. The talk focuses on recurring failures such as invented schemas, excessive context, and repeated tool strategies.

The published model card reports 59.70% on the Snorkel Finance Benchmark for rLLM-FinQA-4B, compared with 51.37% for Qwen3-235B-A22B in the same table. The creator reports training cost below $500. These results support specialization on this benchmark, but they do not show that a 4B model generally outperforms larger models across financial work.

Inspect the published rLLM-FinQA-4B model card for the reported scores, base model, tools, and training notes. Review the open-source rLLM repository before adapting the training and evaluation pipeline.

What you will learn

  • A narrow tool environment can teach a small model repeatable behaviors that a general benchmark does not measure.
  • Verified tasks and executable answers make the outcome reward and judge disagreements easier to inspect.
  • Tool traces reveal failures such as invented table names, context flooding, and repeated unsuccessful queries.
  • Transfer tests and general-capability checks matter because one in-domain score cannot establish safe performance on adjacent work.

How to apply this safely

  1. Define one bounded task with a fixed tool surface and collect reviewed examples from permitted source data.
  2. Build a held-out evaluation set before training and preserve the baseline outputs, tool calls, latency, and cost.
  3. Start with deterministic result checks where possible, then inspect reward hacking and failed trajectories manually.
  4. Compare the trained model with the base model and a larger reference on in-domain, transfer, and out-of-scope cases before any financial use.

Important limitations

  • The benchmark, model, and cost claims come from the project team and model card. Learnetto has not reproduced the training run or audited all benchmark examples.
  • Financial question answering can affect material decisions. The talk does not establish regulatory suitability, current-data coverage, or reliability for unsupported filings and tools.

Sources to check

  • rLLM-FinQA-4B model card Primary project record for the model, tools, reported benchmark scores, and license.
  • rLLM Apache-2.0 repository for agent rollouts, evaluation, rewards, and training backends.

Continue learning on Learnetto

AI evals guide

Separate deterministic checks, model judges, and production monitoring.

Best LLM courses

Study fine-tuning, inference, and model limitations before training.