← Benchmarks 🕐 19 min read
Benchmarks

StateBench Finance Replay Model Candidate Registry

This registry turns the finance/quant model landscape into executable

Source ledger: sources/21-benchmarks/statebench-model-candidate-registry-2026-raw.md

This registry turns the finance/quant model landscape into executable StateBench lanes. It is not a leaderboard. A candidate enters the registry when there is a verified artifact, model card, primary paper, vendor engineering post, or already-ingested source that tells us exactly what to run and what cannot be assumed.

Machine-readable candidate manifest: statebench/suites/finance-model-runtime-candidates-v0.json.

Execution queue: statebench/suites/finance-model-runtime-bakeoff-v0.md and statebench/suites/finance-model-runtime-bakeoff-v0.json.

Operating Rule

Evaluate artifacts, not brand names. A base HF checkpoint, GGUF quant, MLX conversion, API route, SageMaker endpoint, and Neuron-compiled artifact are separate candidates. A model that works under CUDA or vLLM is not automatically compatible with MLX, llama.cpp, ROCm, or AWS Neuron.

Qwen2.5-era artifacts are allowed only as specialist controls, especially for Snowflake/Text2SQL where the current runnable specialist lineage still matters. They are not defaults for general StateBench reasoning.

Ollama is excluded. llama.cpp/GGUF remains in scope.

Candidate Lanes

Priority Lane Candidate Artifact Status StateBench Use Preflight / Blocker
1 Frontier control Best available frontier API model API route Ceiling for all 118 finance replay tasks Cost cap, deterministic params, no live-web score conflation
2 Current Qwen open base Qwen/Qwen3.6-35B-A3B HF safetensors, Apache-2.0; vLLM/SGLang/Transformers documented; community quantizations exist Long-context source review, repo-level agent tasks, PDF/chart/OCR, finance replay patch tasks Verify exact HF revision; record thinking mode; score each GGUF/MLX/API artifact separately; Neuron unproven
3 Current Gemma open base google/gemma-4-31B plus smaller Gemma 4 variants HF safetensors, Apache-2.0; vLLM/SGLang/Transformers documented Independent open-family comparison for finance document, table, and synthesis tasks Verify exact model size; check local memory/runtime; compare dense vs MoE/smaller variants separately
3A Current Gemma finance adapter Srx7703/gemma-4-31b-financial-adapter on google/gemma-4-31B-it PEFT LoRA adapter, Gemma terms; model card plus GitHub companion repo SEC filing QA, source-grounded earnings-recap synthesis, adapter-vs-base StateBench comparison Tiny n=20 BERTScore eval only; adapter-only artifact; no merged/GGUF/MLX/Neuron artifact verified; companion agent uses API models today
3B Regional Gemma2 Turkish finance adapter Sengil/turkish-gemma-9b-finance-sft on ytu-ce-cosmos/Turkish-Gemma-9b-T1 Adapter-only safetensors, revision 7b2a6ade37c2f8065938e8f5718e7089d62d539b; base Gemma2ForCausalLM, base license metadata gemma; card uses Unsloth and Turkish finance SFT datasets Turkish finance-language Gemma-family control through finance-regional-language-v0, compared against Mihenk/Qwen3.6 Turkish lane Base-license review, adapter load, and output hygiene required; no merged/GGUF/MLX/Neuron portability proof; no Ollama
4 Liquid runtime / nano lane LFM2.5 text, VL, audio, Extract, RAG, Tool, Transcript, ColBERT, maximaverick/LFM2.5-1.2B-Financial-Analyst-Thinking, nexus-syntegra/Nexus-TinyFunction-1.2B-v2.0 HF/GGUF/MLX/ONNX availability varies by model; LFM license; Financial Analyst has Apache-2.0 HF metadata and GGUF files; TinyFunction has Apache-2.0 plus LFM license file Low-latency extraction, source-pack QA, podcast transcript cleanup, tool routing, ColBERT reranking, regional A-share/CFA-style explanation via finance-regional-language-v0, finance tool-router controls via finance-tool-routing-v0 Not a broad finance reasoning default; Financial Analyst now has regional fixture coverage but still needs runtime smoke; TinyFunction now has tool-schema fixture coverage but still needs local runtime/BFCL or router eval
5 Finance specialist rLLM-FinQA-4B, ODA-Fin, TheFinAI/Fin-o1-8B, khazarai/Fino1-4B, TraceAlchemy Gemma finance, Fin-R1 controls, Kuvera, CPA-Qwen3, FinGuard, Dev9124/qwen3-finance-model Mixed HF/paper artifacts Filing/table QA, financial reasoning, sentiment, personal-finance suitability, accounting/compliance, compliance guards, source-grounded wiki maintenance Verify license, artifact format, exact data lineage, jurisdiction, and no benchmark-only overclaim
5A Local finance GGUF specialist mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF Apache-2.0 community GGUF quants of a Qwen3-8B finance merge llama.cpp finance-reasoning control, local replay/accounting smoke tasks, quantization comparison Community quant; exact quant file and upstream revision required; no Ollama; no MLX/Neuron portability assumption
5B Regional Qwen3.6 finance SLM AlicanKiraz0/Mihenk-LLM-v2-35B-A3B-Turkish-Financial-Model Apache-2.0 merged HF safetensors on Qwen/Qwen3.6-35B-A3B; Turkish/English, BIST, crypto, risk tags Turkish/BIST finance-language QA and Qwen3.6 finance fine-tune comparison through finance-regional-language-v0 Regional scope only; no U.S. SEC/accounting/institutional-quant transfer; HF artifact scored separately from GGUF/MLX/Neuron
5B-GGUF Regional Qwen3.6 finance GGUF emircansevdi/Mihenk-LLM-v2-35B-A3B-Turkish-Financial-Model-GGUF Apache-2.0 community GGUF conversion, revision f55fff76ca6ce3f262a4e853a2dfb33a635e0914; Q4/Q5/Q6/Q8/BF16 GGUF files; qwen35moe, 34.66B params, 262k context Local llama.cpp Turkish/BIST finance-language control and quantization comparison through finance-regional-language-v0 Community conversion; exact GGUF file and llama.cpp smoke required; no MLX/CUDA/ROCm/SageMaker/Bedrock/Neuron portability proof; no Ollama
5C SEC extraction GGUF TheTokenFactory/gemma-4-E2B-sec-extraction-GGUF-v3 Gemma-license GGUF/llama.cpp artifact on unsloth/gemma-4-E2B-it; card reports unverified extraction metrics Local SEC structured JSON extraction and filing-snippet source-card extraction tests Narrow extraction model; card metrics unverified; exact GGUF/parser/citation reconciliation required
5D Bookkeeping/accounting GGUF Myyyyyyyyyyyyyy/qwen3-14b-bookkeeper-gguf Apache-2.0 Qwen3-14B GGUF trained with Unsloth/QLoRA; accounting/bookkeeping plus reasoning datasets Local bookkeeping/accounting workflow control: transaction categorization, journal entries, reconciliation, accounting guardrails Model card says first attempt/work in progress; synthetic/exam data does not prove GAAP, tax, audit, fraud, or signoff reliability
5E Personal-finance advisor SLM Akhil-Theerthala/Kuvera-8B-qwen3-v0.2.1 MIT Qwen3-8B personal-finance specialist with dataset and paper source card Suitability, personal-finance arithmetic, refusal/human-review, and advisor-support checks in finance-advisor-portfolio-v0 Not institutional quant research, alpha evidence, regulated advice authority, or portfolio-execution proof; no MLX/GGUF/Neuron portability proof
5F Multilingual finance/economics GGUF vividdream/Qwen-Open-Finance-R-8B-IQ4_NL-GGUF Apache-2.0 imatrix GGUF conversion of DragonLLM/Qwen-Open-Finance-R-8B; English/French/German tags Local multilingual finance/economics QA, source-card reasoning, and llama.cpp finance-reasoning control Community quant; QA tags do not prove citation quality, accounting correctness, quant research, alpha, or portfolio construction
5G Options/GEX LoRA specialist dtarkenton/sprocket-gex-deepseek-r1-distill-qwen14b-lora-paper-exact-final on unsloth/DeepSeek-R1-Distill-Qwen-14B-bnb-4bit PEFT/QLoRA adapter, revision 3485e4d64f38818e6860c79430fd606bf827d741, license: other; card claims 2,027 paper-exact SFT rows from a broader SPY+QQQ GEX/options pipeline Structured Gamma Exposure pattern and 30-day regime classification experiments for options/derivatives fixtures Adapter-only, license review required, partly synthetic/paper-exact data, 32-case smoke eval is not robustness evidence; no trading/advice/alpha/portfolio/pricing/Neuron/GGUF/MLX proof; no Ollama
5H Tiny 10-K filing-language SLM HarryS64/10k-financial-slm MIT PyTorch model.pt + train.py, revision 9f60e3dca26a991d1b751d41c889ef620b6db883; 11.5M GPT decoder-only model; trained on 1,131 SEC 10-K filings from financial companies; self-reported val bpb 1.645 vs. 2.146 general-text control Cheap 10-K language-modeling, compression, unusual-language/anomaly, and filing-prefilter experiments before expensive model review Base next-token model, not a chatbot; no QA, table, accounting, citation, alpha, advice, or portfolio proof; PyTorch/MPS evidence does not imply Transformers, MLX, GGUF, CUDA, ROCm, SageMaker, Bedrock, or Neuron
5I Chinese financial-report metric extraction adapter lmxxf/financial-report-lora-qwen3-14b on Qwen/Qwen3-14B PEFT LoRA adapter, HF revision f17d53403bce36e9c9f3f30699aa6d02267a57b7; Chinese structured JSON metric extraction from research-report paragraphs; license not visible in card/API Chinese broker/research-report metric extraction, schema-checked source cards, and upstream factor-discovery metric candidates Adapter-only; 460-row training data and no independent benchmark; license/base revision/adapter load/JSON schema preflight required; no GGUF/MLX/Neuron/backend portability proof; no Ollama
6 Sentiment specialist FinSenti-Qwen3.5, ModernFinBERT, FinBERT controls, YorkFr/financial-sentiment-qwen3-v2, p988744/eland-sentiment-zh-vllm, other Qwen3 sentiment variants Mixed HF artifacts; Eland is Apache-2.0 Qwen3-4B Traditional Chinese/Taiwan-market sentiment model optimized for vLLM Headline, earnings-snippet, analyst-commentary, entity sentiment, opinion sentiment, and market-news routing Score against held-out finance sentiment/routing fixtures; Eland needs Traditional Chinese/Taiwan fixture, explicit system prompt, exact revision capture, and no alpha/portfolio inference
6C Earnings-call evasion specialist FutureMa/Eva-4B-V2 / EvasionBench Apache-2.0 Qwen3-4B-Instruct-2507 derivative; HF BF16 safetensors; arXiv 2601.09142 Earnings-call analyst-question / management-answer directness, evasiveness, and disclosure-quality triage Not sentiment, alpha, advice, management-quality truth, or compliance signoff; label-source/judge bias and human-review caveats required
6B Tiny local sentiment GGUF mad-lab-ai/qwen3-1.7b-sentiment-gguf Apache-2.0 Qwen3-1.7B Q8_0 GGUF, finance/stock-market sentiment QLoRA Cost-floor local sentiment labeler for finance-sentiment-routing-v0 Directional sentiment does not imply alpha; RSI/MACD/EMA-aligned labels require leakage checks; label score does not imply route/rationale/abstention
6A Regional and narrow finance SLMs iravikr/qwen3-0.6b-finance-india, KBTG-Labs/THaLLE-0.2-ThaiLLM-8B-fa, Ayansk11/qwen3-4b-financial-sentiment-grpo, sweatSmile/Qwen3-4B-Instruct-FinanceQA HF model-card metadata verified; mixed merged, mergekit, GGUF, and adapter artifacts India budget/policy QA and Thai/English finance language through finance-regional-language-v0, small sentiment routing, historical FinanceQA adapter controls Use only when jurisdiction/language/task fit matches; exact metadata matters because one Qwen3-named artifact reports Qwen2.5 base lineage
7 Finance multimodal VLM AITRADER/Amsi-fin-o1.5-mxfp8-MLX MLX conversion of finance Qwen3.5-VL parent Financial chart/document/screenshot understanding, OCR repair, multimodal source triage MLX smoke test, parent revision, memory, vision prompt contract; not a text replay default
7A Japanese finance multimodal VLM Yana/compass-vlm Apache-2.0 HF Transformers/Safetensors VLM, revision fa4fe4ac44ab4075f8b749e7266ab494038c6f14; 8.59B BF16 params; Japanese/English; SigLIP/LLM-JP/Qwen3 teacher stack; finance/document-understanding tags EDINET-style Japanese financial statement images/PDFs, chart/table/OCR repair, and multimodal source triage VLM-specific artifact; not U.S. SEC/English text replay/alpha/advice proof; HF safetensors does not imply MLX, GGUF, llama.cpp, SageMaker, Bedrock, or Neuron
8 Retrieval/reranking Qwen3 Embedding/Reranker, FinE5, BGE-M3/BGE reranker, Snowflake Arctic Embed, LFM2-ColBERT, PeterAM4/Qwen3-Embedding-0.6B-GGUF, souflex56/qanchor-reranker-qwen3-0.6b-merged Mixed open/gated/vendor artifacts; QAnchor is Apache-2.0 HF Transformers sequence-classification reranker for Chinese A-share annual-report chunks Source acquisition, citation localization, FinRetrieval-style tasks, Chinese financial-document reranking, retrieval-cost-floor tests Requires retrieval gold set, hybrid BM25 baseline, reranker trace, source/page precision; QAnchor additionally needs qwen3_template formatting and same-document scope checks
8H Finance embedding Gemma shm2290/finance-embeddings-gemma-300m-v2 HF Safetensors, 303M F32 encoder fine-tune on google/embeddinggemma-300m; source fixture sa-109 Small finance-domain embedding control for source localization, finance wiki/static-index retrieval, and RAG corpus search License inherits from base EmbeddingGemma; no independent FinMTEB/StateBench score; no GGUF/MLX/Neuron proof
8I Closed insurance-domain reference AntGroup Finix-S1 Closed/proprietary CUFEInse benchmark signal; source fixture sa-133; no public artifact verified Reference ceiling for insurance-domain fixtures: policy/product knowledge, business-process understanding, compliance, agent application, and logical rigor Not runnable; no checkpoint/API/license/architecture/runtime proof; not institutional quant research, investment alpha, or backend compatibility evidence
8A Cautionary/private corpus controls Mikkkkoooo/qwen35-4b-private-analyst-full-corpus BF16 safetensors plus Q4_K_M GGUF; license other; private research-report corpus Analyst-style source synthesis and contamination/overfitting tests only License and data lineage block promotion; no independent benchmark
8B Web3/DeFi finance SLM DMindAI/DMind-3-mini Apache-2.0 Qwen3.5-4B derivative Crypto/Web3/DeFi source tasks and smart-contract/security-adjacent finance tasks Not an institutional asset-management model; separate Web3 fixture required
8C Tiny banking voice/tool SLM distil-labs/distil-lfm25-voice-assistant 354M BF16 LiquidAI/LFM2.5-350M banking voice-assistant fine-tune, Liquid open license terms Banking voice-agent, tool-routing, and customer-service workflow fixtures Not investment research, trading, compliance, or alpha competence; license/tool-schema fixture required
8D Browser/source-discovery SLM microsoft/Fara-7B MIT open-weight HF artifact; Qwen2.5-VL-based 7B multimodal CUA with 128K context and vLLM/SGLang/Transformers routes Live source-discovery eval: find official pages, PDFs, reports, podcasts, transcripts, and source provenance Not finance reasoning; no MLX/GGUF/Neuron/SageMaker portability assumption
8D-2 Browser/source-discovery SLM Fara1.5-9B on Microsoft Foundry MIT Foundry Version 1 API artifact; Qwen3.5-9B browser CUA with 262K context-window claim and MagenticLite harness path Current-family live source-discovery eval candidate for official pages, PDFs, reports, podcasts, transcripts, and provenance capture Foundry/API route is separate from HF/GGUF/MLX/Neuron artifacts; not finance advice, allocation, alpha, filing analysis, or autonomous high-stakes browser operation
8E Financial time-series / market models NeoQuasar/Kronos-base, Vincent05R/FinCast, and seunghan96/FinSTaR Kronos HF safetensors, MIT license, 102.3M params, OHLCV/K-line tokenizer and 512 context; FinCast source card sa-094 records Apache-2.0 HF/GitHub artifacts, 1B sparse-MoE decoder-only TSFM design, and 20B+ time points; FinSTaR source card sa-096 records public GitHub code, FinTSR-Bench, Compute-in-CoT, Scenario-Aware CoT, and TimeOmni/Qwen2.5 LoRA training Financial K-line forecasting, volatility forecasting, synthetic K-line generation, quantile calibration, variable-frequency forecasting, deterministic price assessment, stochastic prediction reasoning, and market-model baseline comparison in finance-time-series-market-v0 Not chat/source-acquisition models or alpha engines; FinSTaR is code/training-recipe evidence until checkpoint/adapters are verified; requires exact revisions, locked OHLCV/FinTSR fixtures, leakage policy, classical/TSFM baselines, transaction costs, slippage, and runtime smoke proof
8F Financial relation-graph probing Relational Probing with Qwen3 SLM backbones Paper-first arXiv source; Qwen3 0.6B/1.7B/4B upstream SLM experiments; no checkpoint/code artifact verified Future market-graph fixture: induce financial-entity relation graphs from text and feed downstream stock-trend/event prediction models Not a chat model, not a direct trading model, and not runnable yet; requires code/data, point-in-time text, leakage policy, co-occurrence baseline, downstream label contract, and backend/runtime proof
8G Trading/factor reasoning RL Trading-R1 / Alpha-R1 / Trade-R1 Paper-first arXiv source cluster with local PDFs and Chandra OCR; no runnable model artifact verified here Structured thesis generation, context-aware alpha screening, stochastic-reward and reward-hacking guardrails for systematic quant factor research Not live-alpha proof and not deployment-ready; requires code/model/license/revision/runtime verification, point-in-time data, transaction costs, capacity, baselines, and cross-regime reproduction
9 SQL specialist Arctic-Text2SQL-R2 direction; R1/R1.5 runnable controls where available R2 verified as Snowflake engineering post; R1 HF artifact exists Snowflake/Postgres/custom-schema NL2SQL harness R2 open checkpoint not verified in this pass; run only in SQL harness, not patch runner
10 AWS deployment practice SageMaker JumpStart Qwen, SageMaker/Bedrock endpoints, OpenAI-compatible endpoints on AWS AWS announced Qwen3-Coder-Next, Qwen3-30B-A3B, Qwen3-Coder-30B-A3B, Qwen3.5-4B in JumpStart Deployment reps, cost/latency, IAM/logging, endpoint discipline SageMaker availability does not imply Neuron compatibility
11 Neuron cost lane Any above artifact after compile preflight neuron_compiled only after proof Inferentia/Trainium cost reduction experiments Architecture/operator/context/quantization compile gate plus smoke test
12 Proprietary Qwen agent reference Qwen3.7-Max Press/near-primary API signal; no open HF artifact verified API-only long-horizon agent reference for source discovery and tool workflows Capture official docs/API route before scoring; not a local or finance-specialist candidate

Current Preflight Update

The dedicated frontier/API control for finance-vendor-product-signal-v0 now has a small-suite cost estimate at statebench/results/finance-vendor-product-signal-v0/frontier-api-cost-estimate-20260605.json. The exact 12-task backend prompt is 904 input tokens including the system instruction, with a 4,096 output-token cap. A conservative high-cap pricing profile at $15/M input and $75/M output estimates $0.320760 before retries and records a recommended $1.00 cap.

This is preflight evidence only. The control is still not approved for a scored run until an explicit cost_cap and a concrete endpoint/API route are recorded. It also remains API-only evidence and does not imply MLX, GGUF, CUDA, ROCm, SageMaker, Bedrock, or Neuron portability.

sprocket-gex-deepseek-r1-qwen14b-lora is now a watchlist-only options/GEX adapter candidate. Its preflight record pins the adapter revision and artifact files, but scored use is blocked until base-license review, adapter-load smoke, a structured GEX schema fixture, and output-hygiene checks are recorded. Treat it as a derivatives/options schema classifier, not a finance-reasoning, trading-action, alpha, portfolio, or risk-management model.

First Real Run Order

Use the machine-readable candidate manifest before starting a run so artifact format, backend, required preflight, and negative compatibility constraints are recorded consistently.

Command generation:

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --emit-suite-commands finance-replay-v0

Those emitted commands are dry runs by default and should parse without model access. Add --scored only after exact revision, hardware, context length, and backend constraints are replaced with verified values.

For batch review, add --emit-batch-script. Dry-run batch commands still write StateBench result directories if executed, so use generated commands with --list-tasks when only checking suite selection.

Target-suite validation:

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --suite-target-report

Readiness report:

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --readiness-report

Artifact preflight template:

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --preflight-template \
  --artifact-id qwen3-6-35b-a3b-hf

Write and validate preflight records:

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --manifest statebench/suites/finance-model-runtime-candidates-v0.json \
  --preflight-template \
  --write-preflight-template \
  --artifact-id qwen3-6-35b-a3b-hf

scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
  --validate-preflight-record statebench/preflights/qwen3-6-35b-a3b-hf-preflight.json

The validator enforces statebench/schemas/preflight_record.schema.json before checking whether all required preflight items have evidence. Scored finance-replay-v0 suite runs and finance-retrieval-sql-harness-v0 backend calls require --preflight-record; dry-runs, --list-tasks, fixture smoke checks, prompt export, Postgres preflight, and scoring an existing answer artifact remain available without approval.

  1. Run a frontier API control against all 118 finance-replay-v0 tasks.
  2. Run Qwen/Qwen3.6-35B-A3B through an OpenAI-compatible vLLM or SGLang endpoint, with exact HF revision, 262K effective context unless extended context is deliberately enabled, and a thinking-mode note.
  3. Run google/gemma-4-31B or the best Gemma 4 artifact that fits available hardware; record whether the run is dense, MoE, API, GGUF, or MLX.
  4. Run one Liquid task-specialist artifact on the subset it should plausibly improve: source-pack QA, transcript cleanup, extract/PII, RAG, tool routing, or reranking. Do not force it into general finance reasoning.
  5. Run one finance specialist only after its license and artifact are verified. rLLM-FinQA-4B, ODA-Fin-style Qwen3 finance post-training, and TheFinAI/Fin-o1-8B are higher signal than old BloombergGPT/PIXIU references because they are closer to local StateBench experiments. khazarai/Fino1-4B is now the cheaper Qwen3-4B FinQA-style control in the same specialist lane, but it remains blocked on license-conflict and runtime preflight checks. 5a. Do not leave new static-control suites without candidate artifacts. The manifest now targets finance-treasury-payments-liquidity-v0, finance-agentic-factor-discovery-v0, and finance-portfolio-risk-construction-v0 with frontier/API ceiling control, Qwen3.6 open-base comparison, Gemma 4 where the task is factor/portfolio reasoning, and Liquid Financial Analyst only as a treasury/payments tiny-model cost-floor control. These are planning rows, not model scores. 5b. Treat Trading-R1 / Alpha-R1 / Trade-R1 as a separate paper-first trading/factor RL lane. They are highly relevant to systematic quant research because they target structured trading theses, factor screening, noisy market rewards, and reward hacking, but they should not enter a scored runtime queue until code, weights, license, exact revision, backend, transaction-cost assumptions, point-in-time data, and reproducibility evidence are recorded.
  6. Use finance-sentiment-routing-v0 before using large chat models for headline or earnings-snippet classification. The 15-task held-out fixture and static gold-control scorecard now exist. FinSenti-Qwen3.5-9B is the current Qwen-family candidate to compare against FinBERT/ModernFinBERT and tiny Qwen3-0.6B sentiment controls such as YorkFr/financial-sentiment-qwen3-v2, but it is not runnable in this repo until the Qwen3.5 image-text-to-text adapter, FinSenti tag parser, or llama.cpp/GGUF backend is proven. 6a. Add an earnings-call evasion sublane. Eva-4B-V2 is a Qwen3-4B-derived specialist for analyst-question / management-answer directness and evasiveness. It belongs in transcript/disclosure-quality triage and earnings-call feature extraction, not in generic sentiment, alpha, investment advice, or portfolio construction. Any score must record label-source bias, LLM-as-judge risk, intermediate-class subjectivity, data-vintage drift, and human-review requirements.
  7. Build the accounting/compliance guard lane separately from broad finance reasoning. CPA-Qwen3 belongs on accounting/tax/audit boundary tasks, while FinGuard belongs on policy-grounded verifier tasks after jurisdiction and policy corpus are defined.
  8. Build the multimodal finance lane separately from the text patch runner. AITRADER/Amsi-fin-o1.5-mxfp8-MLX is useful for chart/OCR/screenshot source tasks only after an mlx-vlm smoke test proves it runs locally.
  9. Build the retrieval lane separately from the patch runner. Start with Qwen3 Embedding plus Qwen3 Reranker against FinRetrieval-style and internal source-localization tasks, then compare finance-specific/gated encoders such as FinE5 if licensing permits.
  10. Build the SQL lane separately from the patch runner. Treat Snowflake Arctic-Text2SQL-R2 as the current Snowflake-specialist recipe to test, but do not score vendor-post claims as internal performance.
  11. Only after base/API/local runs identify repeatable failure modes, compare Unsloth, LLaMA-Factory, AWS native, and MLX fine-tuning paths.

Required Manifest Metadata

Every run must fill these fields in statebench/results/<suite_run_id>/manifest.json:

  • model
  • weights_revision
  • artifact_id
  • adapter_id
  • merge_state
  • artifact_format
  • artifact_precision
  • backend_id
  • backend_hardware
  • backend_constraints
  • context_length

For backend constraints, record negative knowledge explicitly, for example:

  • neuron compatibility unproven until operator and compilation preflight passes
  • mlx conversion not tested
  • gguf quantization is community artifact, not upstream release
  • api-only proprietary route
  • license review required before commercial use

Why This Matters For The Finance Goal

The repo already has broad finance/quant coverage. The remaining risk is selection drift: choosing models because they are fashionable, locally convenient, or vendor-promoted instead of because they pass the finance tasks we actually need.

This registry prevents three common mistakes:

  1. Treating Qwen, Gemma, or Liquid family names as scores.
  2. Letting SQL/retrieval/source-discovery performance leak into the general patch-runner benchmark.
  3. Assuming deployment portability across CUDA, MLX, llama.cpp, AMD, and Neuron without explicit artifact evidence.

Current Interpretation

As of May 29, 2026:

  • Qwen/Qwen3.6-35B-A3B is the default current open Qwen-family base candidate for StateBench finance replay. It is current enough to avoid the stale Qwen2.5 trap and has explicit long-context and agentic-coding evidence.
  • Gemma 4 is the right open-family counterweight to Qwen, especially for multimodal and long-context document tasks.
  • LiquidAI is most interesting as a family of small task specialists and portable runtime artifacts, not as a finance alpha/reasoning model.
  • Arctic-Text2SQL-R2 is the most important new SQL signal because it argues that compact, domain/dialect-trained models can beat frontier systems on enterprise SQL. That belongs in a Snowflake/Postgres harness, not the general finance replay patch runner.
  • Dev9124/qwen3-finance-model is a current Qwen3 finance-instruction watchlist artifact with Apache-2.0 model-card licensing and vLLM/SGLang runtime snippets. It is useful as a broad fine-tune control, but model-card evidence alone is weaker than ODA-Fin/rLLM-style paper-backed specialists.
  • TheFinAI/Fin-o1-8B is now tracked as a Qwen3-8B finance-reasoning candidate. The current HF card supersedes the stale local Llama-lineage note; score it as a concrete Qwen3 artifact only after exact revision and runtime proof.
  • khazarai/Fino1-4B is now tracked as a separate Qwen3-4B finance-reasoning watchlist candidate, not a smaller alias for TheFinAI/Fin-o1-8B. The captured HF revision is 15c0d969d0bd2aa6515e0e315b5aecc063c2d7ee, architecture is Qwen3ForCausalLM, base metadata is unsloth/Qwen3-4B, and the model has 4,022,468,096 BF16 safetensors parameters. Its dataset points to TheFinAI/Fino1_Reasoning_Path_FinQA at revision 0316559663c46e50d90e75f6e14a6d8f679cdba8 under cc-by-4.0. Block production/commercial assumptions until the model license conflict is resolved: HF metadata says Apache-2.0 while the README body says MIT. Treat it as a cheap FinQA-style numerical/table-QA control only; no GGUF, MLX, ROCm, SageMaker, Bedrock, or Neuron portability is proven.
  • CPA-Qwen3 and FinGuard widen the domain-specialist lane from investment reasoning into accounting/compliance and policy-grounded guardrails. Their scores should live on accounting/compliance fixtures, not generic replay. The first finance-accounting-compliance-v0 seed fixture now exists with 10 guardrail tasks and a static 10/10 answer-artifact control; model/backend scores still require preflight approval.
  • AITRADER/Amsi-fin-o1.5-mxfp8-MLX adds a current MLX finance VLM candidate for charts, OCR, screenshots, and multimodal source triage. It is not proof that MLX should replace other runtimes, and it is not a text-only benchmark default.
  • YorkFr/financial-sentiment-qwen3-v2 is a tiny MIT-licensed Qwen3-0.6B sentiment artifact. It belongs in a held-out sentiment/routing lane, not the general finance replay patch runner. The fixture and Transformers-local adapter now exist. The first local run is negative evidence: 0/15 accepted, despite partial label signal, because route/rationale controls failed.
  • Regional and narrow finance SLMs are now tracked as a separate watchlist: iravikr/qwen3-0.6b-finance-india for India budget/policy QA, KBTG-Labs/THaLLE-0.2-ThaiLLM-8B-fa for Thai/English finance-language coverage, and Ayansk11/qwen3-4b-financial-sentiment-grpo as a smaller Qwen3 sentiment/GGUF candidate. sweatSmile/Qwen3-4B-Instruct-FinanceQA is a useful cautionary control because its Hugging Face metadata reports a Qwen2.5 base despite the repository name.
  • The May 31 HF watchlist pass adds a more concrete local runtime candidate: mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF, revision 29d21ffe58d0e9c7f86136fe7f9eaab3eb0d1332, as a llama.cpp/GGUF finance reasoning control. It should be scored as its exact quant file, not as the upstream RinKana merge. The same pass adds PeterAM4/Qwen3-Embedding-0.6B-GGUF, revision bb661aeeeafa4ff7303b5f5e1e80255efd6a513c, as a tiny local embedding candidate with finance-related imatrix/data tags.
  • The June 5 HF watchlist follow-up adds shm2290/finance-embeddings-gemma-300m-v2 as a finance-specific EmbeddingGemma 300M control. It is useful precisely because it is small and retrieval-only: test it on source localization, wiki/static-index retrieval, SEC/investor-document paragraph retrieval, and podcast/source-card metadata search. Do not use it as a finance reasoner, source-discovery agent, accounting model, or alpha signal. Its license inherits from google/embeddinggemma-300m, so commercial and redistribution use need base license review before promotion.
  • The June 6 HF watchlist follow-up adds souflex56/qanchor-reranker-qwen3-0.6b-merged as a Chinese financial-document reranker candidate. It is useful because it targets A-share annual reports and related filings with a current Qwen3 sequence-classification reranker, but it is not a generator or broad finance model. Score it only after a pinned Chinese financial-document retrieval fixture, first-stage BM25/embedding/RRF baseline, preserved qwen3_template formatting, same-document scope check, and reranker trace schema are in place.
  • The June 6 regional-sentiment follow-up adds p988744/eland-sentiment-zh-vllm as a Qwen3-4B Traditional Chinese/Taiwan financial-sentiment routing candidate. It is useful for Taiwan stock-market headlines, entity sentiment, and opinion sentiment, but it is not a U.S. SEC, accounting, quant-research, investment-advice, portfolio, trading, or alpha model. Score only after exact HF revisions are captured, the required system prompt is preserved, a held-out Traditional Chinese/Taiwan fixture exists, and vLLM/PEFT/GGUF artifacts are separated. Do not use Ollama.
  • The June 6 structured-extraction follow-up adds lmxxf/financial-report-lora-qwen3-14b as a Chinese financial research-report metric-extraction adapter on Qwen/Qwen3-14B. It is useful because it targets schema-checked extraction of deeply analyzed metrics from Chinese report paragraphs, a better fit for source cards and factor-discovery bootstrapping than generic sentiment. It is blocked on license visibility, base revision, PEFT adapter-load smoke, held-out Chinese report metric spans, JSON-schema validation, and output-hygiene checks. Do not treat it as a broad finance reasoner, accounting authority, alpha model, or runtime portability proof.
  • Mikkkkoooo/qwen35-4b-private-analyst-full-corpus and DMindAI/DMind-3-mini are recorded as watchlist controls only. The former is a private-corpus analyst-style fine-tune with license other; the latter is a Web3/DeFi finance model, not a U.S. asset-management or SEC-filing model.
  • The May 31 follow-up pass adds three narrow candidates. Mihenk is a Qwen3.6-35B-A3B Turkish/BIST finance model now covered by the finance-regional-language-v0 scope-control fixture. Distil/LFM2.5 voice assistant is a 354M banking tool-calling candidate for voice/tool-routing fixtures, not quant research. TokenFactory Gemma SEC extraction v3 is a GGUF/llama.cpp structured-extraction candidate with unverified card metrics and Gemma-license constraints.
  • The June 5 follow-up records a separate community GGUF conversion of Mihenk. This improves the local llama.cpp runtime lane but does not change the upstream HF/BF16 evidence. The GGUF artifact is blocked until an exact file, prompt-template compatibility, and llama.cpp smoke test are recorded.
  • The June 5 Gemma follow-up records Sengil/turkish-gemma-9b-finance-sft as a regional Turkish finance-language Gemma2 adapter. It is useful as a family-control against the Qwen/Mihenk Turkish lane, but it remains blocked on base-license review, adapter-load smoke, and visible reasoning/output-hygiene checks.
  • The second May 31 follow-up pass adds three more executable local controls. Myyyyyyyyyyyyyy/qwen3-14b-bookkeeper-gguf fills a bookkeeping/accounting workflow gap with Qwen3-14B GGUF and explicit work-in-progress caveats. mad-lab-ai/qwen3-1.7b-sentiment-gguf is a tiny local sentiment-cost-floor candidate for label/routing fixtures, with leakage and alpha-overclaim guardrails. vividdream/Qwen-Open-Finance-R-8B-IQ4_NL-GGUF gives a multilingual finance/economics llama.cpp control that must be scored separately from upstream DragonLLM weights and other quantizations.
  • neoyipeng/ModernFinBERT-base and ProsusAI/finbert establish label-only encoder baselines on the same fixture. ModernFinBERT currently leads the label slice (0.733 vs. 0.667 for ProsusAI and YorkFr), but both encoders fail full routing acceptance because they do not emit route/rationale policy.
  • AssetOpsBench QLoRA tool-knowledge evidence strengthens the fine-tuning gate: small models may internalize stable tool catalogs, but only after StateBench shows repeated tool/planning failures and after catastrophic forgetting is measured.
  • Qwen3.7-Max is worth tracking as a proprietary long-horizon agent reference, but it is not an open/local candidate and should not displace Qwen3.6 for the local artifact lane without verified open weights.
  • Relational Probing adds a distinct market-graph lane: Qwen3 0.6B/1.7B/4B SLMs are used as upstream hidden-state encoders whose relation head induces financial-entity graphs for downstream stock-trend prediction. This is closer to quant modeling than finance chat, but it is paper-first evidence until code/data/checkpoints are verified and leakage-safe point-in-time fixtures exist.