What this video shows
Dickens describes a financial question-answering agent with table-listing, schema, SQL, and calculator tools. The team starts from Qwen3-4B-Instruct and trains with reinforcement learning on verified questions drawn from SEC 10-K filings. The talk focuses on recurring failures such as invented schemas, excessive context, and repeated tool strategies.
The published model card reports 59.70% on the Snorkel Finance Benchmark for rLLM-FinQA-4B, compared with 51.37% for Qwen3-235B-A22B in the same table. The creator reports training cost below $500. These results support specialization on this benchmark, but they do not show that a 4B model generally outperforms larger models across financial work.
Inspect the published rLLM-FinQA-4B model card for the reported scores, base model, tools, and training notes. Review the open-source rLLM repository before adapting the training and evaluation pipeline.
What you will learn
- A narrow tool environment can teach a small model repeatable behaviors that a general benchmark does not measure.
- Verified tasks and executable answers make the outcome reward and judge disagreements easier to inspect.
- Tool traces reveal failures such as invented table names, context flooding, and repeated unsuccessful queries.
- Transfer tests and general-capability checks matter because one in-domain score cannot establish safe performance on adjacent work.
How to apply this safely
- Define one bounded task with a fixed tool surface and collect reviewed examples from permitted source data.
- Build a held-out evaluation set before training and preserve the baseline outputs, tool calls, latency, and cost.
- Start with deterministic result checks where possible, then inspect reward hacking and failed trajectories manually.
- Compare the trained model with the base model and a larger reference on in-domain, transfer, and out-of-scope cases before any financial use.
Important limitations
- The benchmark, model, and cost claims come from the project team and model card. Learnetto has not reproduced the training run or audited all benchmark examples.
- Financial question answering can affect material decisions. The talk does not establish regulatory suitability, current-data coverage, or reliability for unsupported filings and tools.
Sources to check
- rLLM-FinQA-4B model card Primary project record for the model, tools, reported benchmark scores, and license.
- rLLM Apache-2.0 repository for agent rollouts, evaluation, rewards, and training backends.
Continue learning on Learnetto
Best AI agent evaluation courses
Design datasets, graders, and trace reviews for tool-using agents.
AI evals guide
Separate deterministic checks, model judges, and production monitoring.
Best LLM courses
Study fine-tuning, inference, and model limitations before training.