Source ledger: sources/21-benchmarks/statebench-model-candidate-registry-2026-raw.md
This registry turns the finance/quant model landscape into executable StateBench lanes. It is not a leaderboard. A candidate enters the registry when there is a verified artifact, model card, primary paper, vendor engineering post, or already-ingested source that tells us exactly what to run and what cannot be assumed.
Machine-readable candidate manifest:
statebench/suites/finance-model-runtime-candidates-v0.json.
Execution queue:
statebench/suites/finance-model-runtime-bakeoff-v0.md
and
statebench/suites/finance-model-runtime-bakeoff-v0.json.
Operating Rule
Evaluate artifacts, not brand names. A base HF checkpoint, GGUF quant, MLX conversion, API route, SageMaker endpoint, and Neuron-compiled artifact are separate candidates. A model that works under CUDA or vLLM is not automatically compatible with MLX, llama.cpp, ROCm, or AWS Neuron.
Qwen2.5-era artifacts are allowed only as specialist controls, especially for Snowflake/Text2SQL where the current runnable specialist lineage still matters. They are not defaults for general StateBench reasoning.
Ollama is excluded. llama.cpp/GGUF remains in scope.
Candidate Lanes
| Priority | Lane | Candidate | Artifact Status | StateBench Use | Preflight / Blocker |
|---|---|---|---|---|---|
| 1 | Frontier control | Best available frontier API model | API route | Ceiling for all 118 finance replay tasks | Cost cap, deterministic params, no live-web score conflation |
| 2 | Current Qwen open base | Qwen/Qwen3.6-35B-A3B |
HF safetensors, Apache-2.0; vLLM/SGLang/Transformers documented; community quantizations exist | Long-context source review, repo-level agent tasks, PDF/chart/OCR, finance replay patch tasks | Verify exact HF revision; record thinking mode; score each GGUF/MLX/API artifact separately; Neuron unproven |
| 3 | Current Gemma open base | google/gemma-4-31B plus smaller Gemma 4 variants |
HF safetensors, Apache-2.0; vLLM/SGLang/Transformers documented | Independent open-family comparison for finance document, table, and synthesis tasks | Verify exact model size; check local memory/runtime; compare dense vs MoE/smaller variants separately |
| 3A | Current Gemma finance adapter | Srx7703/gemma-4-31b-financial-adapter on google/gemma-4-31B-it |
PEFT LoRA adapter, Gemma terms; model card plus GitHub companion repo | SEC filing QA, source-grounded earnings-recap synthesis, adapter-vs-base StateBench comparison | Tiny n=20 BERTScore eval only; adapter-only artifact; no merged/GGUF/MLX/Neuron artifact verified; companion agent uses API models today |
| 3B | Regional Gemma2 Turkish finance adapter | Sengil/turkish-gemma-9b-finance-sft on ytu-ce-cosmos/Turkish-Gemma-9b-T1 |
Adapter-only safetensors, revision 7b2a6ade37c2f8065938e8f5718e7089d62d539b; base Gemma2ForCausalLM, base license metadata gemma; card uses Unsloth and Turkish finance SFT datasets |
Turkish finance-language Gemma-family control through finance-regional-language-v0, compared against Mihenk/Qwen3.6 Turkish lane |
Base-license review, adapter load, and output hygiene required; no merged/GGUF/MLX/Neuron portability proof; no Ollama |
| 4 | Liquid runtime / nano lane | LFM2.5 text, VL, audio, Extract, RAG, Tool, Transcript, ColBERT, maximaverick/LFM2.5-1.2B-Financial-Analyst-Thinking, nexus-syntegra/Nexus-TinyFunction-1.2B-v2.0 |
HF/GGUF/MLX/ONNX availability varies by model; LFM license; Financial Analyst has Apache-2.0 HF metadata and GGUF files; TinyFunction has Apache-2.0 plus LFM license file | Low-latency extraction, source-pack QA, podcast transcript cleanup, tool routing, ColBERT reranking, regional A-share/CFA-style explanation via finance-regional-language-v0, finance tool-router controls via finance-tool-routing-v0 |
Not a broad finance reasoning default; Financial Analyst now has regional fixture coverage but still needs runtime smoke; TinyFunction now has tool-schema fixture coverage but still needs local runtime/BFCL or router eval |
| 5 | Finance specialist | rLLM-FinQA-4B, ODA-Fin, TheFinAI/Fin-o1-8B, khazarai/Fino1-4B, TraceAlchemy Gemma finance, Fin-R1 controls, Kuvera, CPA-Qwen3, FinGuard, Dev9124/qwen3-finance-model |
Mixed HF/paper artifacts | Filing/table QA, financial reasoning, sentiment, personal-finance suitability, accounting/compliance, compliance guards, source-grounded wiki maintenance | Verify license, artifact format, exact data lineage, jurisdiction, and no benchmark-only overclaim |
| 5A | Local finance GGUF specialist | mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF |
Apache-2.0 community GGUF quants of a Qwen3-8B finance merge | llama.cpp finance-reasoning control, local replay/accounting smoke tasks, quantization comparison | Community quant; exact quant file and upstream revision required; no Ollama; no MLX/Neuron portability assumption |
| 5B | Regional Qwen3.6 finance SLM | AlicanKiraz0/Mihenk-LLM-v2-35B-A3B-Turkish-Financial-Model |
Apache-2.0 merged HF safetensors on Qwen/Qwen3.6-35B-A3B; Turkish/English, BIST, crypto, risk tags |
Turkish/BIST finance-language QA and Qwen3.6 finance fine-tune comparison through finance-regional-language-v0 |
Regional scope only; no U.S. SEC/accounting/institutional-quant transfer; HF artifact scored separately from GGUF/MLX/Neuron |
| 5B-GGUF | Regional Qwen3.6 finance GGUF | emircansevdi/Mihenk-LLM-v2-35B-A3B-Turkish-Financial-Model-GGUF |
Apache-2.0 community GGUF conversion, revision f55fff76ca6ce3f262a4e853a2dfb33a635e0914; Q4/Q5/Q6/Q8/BF16 GGUF files; qwen35moe, 34.66B params, 262k context |
Local llama.cpp Turkish/BIST finance-language control and quantization comparison through finance-regional-language-v0 |
Community conversion; exact GGUF file and llama.cpp smoke required; no MLX/CUDA/ROCm/SageMaker/Bedrock/Neuron portability proof; no Ollama |
| 5C | SEC extraction GGUF | TheTokenFactory/gemma-4-E2B-sec-extraction-GGUF-v3 |
Gemma-license GGUF/llama.cpp artifact on unsloth/gemma-4-E2B-it; card reports unverified extraction metrics |
Local SEC structured JSON extraction and filing-snippet source-card extraction tests | Narrow extraction model; card metrics unverified; exact GGUF/parser/citation reconciliation required |
| 5D | Bookkeeping/accounting GGUF | Myyyyyyyyyyyyyy/qwen3-14b-bookkeeper-gguf |
Apache-2.0 Qwen3-14B GGUF trained with Unsloth/QLoRA; accounting/bookkeeping plus reasoning datasets | Local bookkeeping/accounting workflow control: transaction categorization, journal entries, reconciliation, accounting guardrails | Model card says first attempt/work in progress; synthetic/exam data does not prove GAAP, tax, audit, fraud, or signoff reliability |
| 5E | Personal-finance advisor SLM | Akhil-Theerthala/Kuvera-8B-qwen3-v0.2.1 |
MIT Qwen3-8B personal-finance specialist with dataset and paper source card | Suitability, personal-finance arithmetic, refusal/human-review, and advisor-support checks in finance-advisor-portfolio-v0 |
Not institutional quant research, alpha evidence, regulated advice authority, or portfolio-execution proof; no MLX/GGUF/Neuron portability proof |
| 5F | Multilingual finance/economics GGUF | vividdream/Qwen-Open-Finance-R-8B-IQ4_NL-GGUF |
Apache-2.0 imatrix GGUF conversion of DragonLLM/Qwen-Open-Finance-R-8B; English/French/German tags |
Local multilingual finance/economics QA, source-card reasoning, and llama.cpp finance-reasoning control | Community quant; QA tags do not prove citation quality, accounting correctness, quant research, alpha, or portfolio construction |
| 5G | Options/GEX LoRA specialist | dtarkenton/sprocket-gex-deepseek-r1-distill-qwen14b-lora-paper-exact-final on unsloth/DeepSeek-R1-Distill-Qwen-14B-bnb-4bit |
PEFT/QLoRA adapter, revision 3485e4d64f38818e6860c79430fd606bf827d741, license: other; card claims 2,027 paper-exact SFT rows from a broader SPY+QQQ GEX/options pipeline |
Structured Gamma Exposure pattern and 30-day regime classification experiments for options/derivatives fixtures | Adapter-only, license review required, partly synthetic/paper-exact data, 32-case smoke eval is not robustness evidence; no trading/advice/alpha/portfolio/pricing/Neuron/GGUF/MLX proof; no Ollama |
| 5H | Tiny 10-K filing-language SLM | HarryS64/10k-financial-slm |
MIT PyTorch model.pt + train.py, revision 9f60e3dca26a991d1b751d41c889ef620b6db883; 11.5M GPT decoder-only model; trained on 1,131 SEC 10-K filings from financial companies; self-reported val bpb 1.645 vs. 2.146 general-text control |
Cheap 10-K language-modeling, compression, unusual-language/anomaly, and filing-prefilter experiments before expensive model review | Base next-token model, not a chatbot; no QA, table, accounting, citation, alpha, advice, or portfolio proof; PyTorch/MPS evidence does not imply Transformers, MLX, GGUF, CUDA, ROCm, SageMaker, Bedrock, or Neuron |
| 5I | Chinese financial-report metric extraction adapter | lmxxf/financial-report-lora-qwen3-14b on Qwen/Qwen3-14B |
PEFT LoRA adapter, HF revision f17d53403bce36e9c9f3f30699aa6d02267a57b7; Chinese structured JSON metric extraction from research-report paragraphs; license not visible in card/API |
Chinese broker/research-report metric extraction, schema-checked source cards, and upstream factor-discovery metric candidates | Adapter-only; 460-row training data and no independent benchmark; license/base revision/adapter load/JSON schema preflight required; no GGUF/MLX/Neuron/backend portability proof; no Ollama |
| 6 | Sentiment specialist | FinSenti-Qwen3.5, ModernFinBERT, FinBERT controls, YorkFr/financial-sentiment-qwen3-v2, p988744/eland-sentiment-zh-vllm, other Qwen3 sentiment variants |
Mixed HF artifacts; Eland is Apache-2.0 Qwen3-4B Traditional Chinese/Taiwan-market sentiment model optimized for vLLM | Headline, earnings-snippet, analyst-commentary, entity sentiment, opinion sentiment, and market-news routing | Score against held-out finance sentiment/routing fixtures; Eland needs Traditional Chinese/Taiwan fixture, explicit system prompt, exact revision capture, and no alpha/portfolio inference |
| 6C | Earnings-call evasion specialist | FutureMa/Eva-4B-V2 / EvasionBench |
Apache-2.0 Qwen3-4B-Instruct-2507 derivative; HF BF16 safetensors; arXiv 2601.09142 | Earnings-call analyst-question / management-answer directness, evasiveness, and disclosure-quality triage | Not sentiment, alpha, advice, management-quality truth, or compliance signoff; label-source/judge bias and human-review caveats required |
| 6B | Tiny local sentiment GGUF | mad-lab-ai/qwen3-1.7b-sentiment-gguf |
Apache-2.0 Qwen3-1.7B Q8_0 GGUF, finance/stock-market sentiment QLoRA | Cost-floor local sentiment labeler for finance-sentiment-routing-v0 |
Directional sentiment does not imply alpha; RSI/MACD/EMA-aligned labels require leakage checks; label score does not imply route/rationale/abstention |
| 6A | Regional and narrow finance SLMs | iravikr/qwen3-0.6b-finance-india, KBTG-Labs/THaLLE-0.2-ThaiLLM-8B-fa, Ayansk11/qwen3-4b-financial-sentiment-grpo, sweatSmile/Qwen3-4B-Instruct-FinanceQA |
HF model-card metadata verified; mixed merged, mergekit, GGUF, and adapter artifacts | India budget/policy QA and Thai/English finance language through finance-regional-language-v0, small sentiment routing, historical FinanceQA adapter controls |
Use only when jurisdiction/language/task fit matches; exact metadata matters because one Qwen3-named artifact reports Qwen2.5 base lineage |
| 7 | Finance multimodal VLM | AITRADER/Amsi-fin-o1.5-mxfp8-MLX |
MLX conversion of finance Qwen3.5-VL parent | Financial chart/document/screenshot understanding, OCR repair, multimodal source triage | MLX smoke test, parent revision, memory, vision prompt contract; not a text replay default |
| 7A | Japanese finance multimodal VLM | Yana/compass-vlm |
Apache-2.0 HF Transformers/Safetensors VLM, revision fa4fe4ac44ab4075f8b749e7266ab494038c6f14; 8.59B BF16 params; Japanese/English; SigLIP/LLM-JP/Qwen3 teacher stack; finance/document-understanding tags |
EDINET-style Japanese financial statement images/PDFs, chart/table/OCR repair, and multimodal source triage | VLM-specific artifact; not U.S. SEC/English text replay/alpha/advice proof; HF safetensors does not imply MLX, GGUF, llama.cpp, SageMaker, Bedrock, or Neuron |
| 8 | Retrieval/reranking | Qwen3 Embedding/Reranker, FinE5, BGE-M3/BGE reranker, Snowflake Arctic Embed, LFM2-ColBERT, PeterAM4/Qwen3-Embedding-0.6B-GGUF, souflex56/qanchor-reranker-qwen3-0.6b-merged |
Mixed open/gated/vendor artifacts; QAnchor is Apache-2.0 HF Transformers sequence-classification reranker for Chinese A-share annual-report chunks | Source acquisition, citation localization, FinRetrieval-style tasks, Chinese financial-document reranking, retrieval-cost-floor tests | Requires retrieval gold set, hybrid BM25 baseline, reranker trace, source/page precision; QAnchor additionally needs qwen3_template formatting and same-document scope checks |
| 8H | Finance embedding Gemma | shm2290/finance-embeddings-gemma-300m-v2 |
HF Safetensors, 303M F32 encoder fine-tune on google/embeddinggemma-300m; source fixture sa-109 |
Small finance-domain embedding control for source localization, finance wiki/static-index retrieval, and RAG corpus search | License inherits from base EmbeddingGemma; no independent FinMTEB/StateBench score; no GGUF/MLX/Neuron proof |
| 8I | Closed insurance-domain reference | AntGroup Finix-S1 |
Closed/proprietary CUFEInse benchmark signal; source fixture sa-133; no public artifact verified |
Reference ceiling for insurance-domain fixtures: policy/product knowledge, business-process understanding, compliance, agent application, and logical rigor | Not runnable; no checkpoint/API/license/architecture/runtime proof; not institutional quant research, investment alpha, or backend compatibility evidence |
| 8A | Cautionary/private corpus controls | Mikkkkoooo/qwen35-4b-private-analyst-full-corpus |
BF16 safetensors plus Q4_K_M GGUF; license other; private research-report corpus |
Analyst-style source synthesis and contamination/overfitting tests only | License and data lineage block promotion; no independent benchmark |
| 8B | Web3/DeFi finance SLM | DMindAI/DMind-3-mini |
Apache-2.0 Qwen3.5-4B derivative | Crypto/Web3/DeFi source tasks and smart-contract/security-adjacent finance tasks | Not an institutional asset-management model; separate Web3 fixture required |
| 8C | Tiny banking voice/tool SLM | distil-labs/distil-lfm25-voice-assistant |
354M BF16 LiquidAI/LFM2.5-350M banking voice-assistant fine-tune, Liquid open license terms |
Banking voice-agent, tool-routing, and customer-service workflow fixtures | Not investment research, trading, compliance, or alpha competence; license/tool-schema fixture required |
| 8D | Browser/source-discovery SLM | microsoft/Fara-7B |
MIT open-weight HF artifact; Qwen2.5-VL-based 7B multimodal CUA with 128K context and vLLM/SGLang/Transformers routes | Live source-discovery eval: find official pages, PDFs, reports, podcasts, transcripts, and source provenance | Not finance reasoning; no MLX/GGUF/Neuron/SageMaker portability assumption |
| 8D-2 | Browser/source-discovery SLM | Fara1.5-9B on Microsoft Foundry |
MIT Foundry Version 1 API artifact; Qwen3.5-9B browser CUA with 262K context-window claim and MagenticLite harness path | Current-family live source-discovery eval candidate for official pages, PDFs, reports, podcasts, transcripts, and provenance capture | Foundry/API route is separate from HF/GGUF/MLX/Neuron artifacts; not finance advice, allocation, alpha, filing analysis, or autonomous high-stakes browser operation |
| 8E | Financial time-series / market models | NeoQuasar/Kronos-base, Vincent05R/FinCast, and seunghan96/FinSTaR |
Kronos HF safetensors, MIT license, 102.3M params, OHLCV/K-line tokenizer and 512 context; FinCast source card sa-094 records Apache-2.0 HF/GitHub artifacts, 1B sparse-MoE decoder-only TSFM design, and 20B+ time points; FinSTaR source card sa-096 records public GitHub code, FinTSR-Bench, Compute-in-CoT, Scenario-Aware CoT, and TimeOmni/Qwen2.5 LoRA training |
Financial K-line forecasting, volatility forecasting, synthetic K-line generation, quantile calibration, variable-frequency forecasting, deterministic price assessment, stochastic prediction reasoning, and market-model baseline comparison in finance-time-series-market-v0 |
Not chat/source-acquisition models or alpha engines; FinSTaR is code/training-recipe evidence until checkpoint/adapters are verified; requires exact revisions, locked OHLCV/FinTSR fixtures, leakage policy, classical/TSFM baselines, transaction costs, slippage, and runtime smoke proof |
| 8F | Financial relation-graph probing | Relational Probing with Qwen3 SLM backbones | Paper-first arXiv source; Qwen3 0.6B/1.7B/4B upstream SLM experiments; no checkpoint/code artifact verified | Future market-graph fixture: induce financial-entity relation graphs from text and feed downstream stock-trend/event prediction models | Not a chat model, not a direct trading model, and not runnable yet; requires code/data, point-in-time text, leakage policy, co-occurrence baseline, downstream label contract, and backend/runtime proof |
| 8G | Trading/factor reasoning RL | Trading-R1 / Alpha-R1 / Trade-R1 | Paper-first arXiv source cluster with local PDFs and Chandra OCR; no runnable model artifact verified here | Structured thesis generation, context-aware alpha screening, stochastic-reward and reward-hacking guardrails for systematic quant factor research | Not live-alpha proof and not deployment-ready; requires code/model/license/revision/runtime verification, point-in-time data, transaction costs, capacity, baselines, and cross-regime reproduction |
| 9 | SQL specialist | Arctic-Text2SQL-R2 direction; R1/R1.5 runnable controls where available | R2 verified as Snowflake engineering post; R1 HF artifact exists | Snowflake/Postgres/custom-schema NL2SQL harness | R2 open checkpoint not verified in this pass; run only in SQL harness, not patch runner |
| 10 | AWS deployment practice | SageMaker JumpStart Qwen, SageMaker/Bedrock endpoints, OpenAI-compatible endpoints on AWS | AWS announced Qwen3-Coder-Next, Qwen3-30B-A3B, Qwen3-Coder-30B-A3B, Qwen3.5-4B in JumpStart | Deployment reps, cost/latency, IAM/logging, endpoint discipline | SageMaker availability does not imply Neuron compatibility |
| 11 | Neuron cost lane | Any above artifact after compile preflight | neuron_compiled only after proof |
Inferentia/Trainium cost reduction experiments | Architecture/operator/context/quantization compile gate plus smoke test |
| 12 | Proprietary Qwen agent reference | Qwen3.7-Max | Press/near-primary API signal; no open HF artifact verified | API-only long-horizon agent reference for source discovery and tool workflows | Capture official docs/API route before scoring; not a local or finance-specialist candidate |
Current Preflight Update
The dedicated frontier/API control for
finance-vendor-product-signal-v0 now has a small-suite cost estimate at
statebench/results/finance-vendor-product-signal-v0/frontier-api-cost-estimate-20260605.json.
The exact 12-task backend prompt is 904 input tokens including the system
instruction, with a 4,096 output-token cap. A conservative high-cap pricing
profile at $15/M input and $75/M output estimates $0.320760 before
retries and records a recommended $1.00 cap.
This is preflight evidence only. The control is still not approved for a scored
run until an explicit cost_cap and a concrete endpoint/API route are
recorded. It also remains API-only evidence and does not imply MLX, GGUF, CUDA,
ROCm, SageMaker, Bedrock, or Neuron portability.
sprocket-gex-deepseek-r1-qwen14b-lora is now a watchlist-only options/GEX
adapter candidate. Its preflight record pins the adapter revision and artifact
files, but scored use is blocked until base-license review, adapter-load smoke,
a structured GEX schema fixture, and output-hygiene checks are recorded. Treat
it as a derivatives/options schema classifier, not a finance-reasoning,
trading-action, alpha, portfolio, or risk-management model.
First Real Run Order
Use the machine-readable candidate manifest before starting a run so artifact format, backend, required preflight, and negative compatibility constraints are recorded consistently.
Command generation:
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--emit-suite-commands finance-replay-v0
Those emitted commands are dry runs by default and should parse without model
access. Add --scored only after exact revision, hardware, context length, and
backend constraints are replaced with verified values.
For batch review, add --emit-batch-script. Dry-run batch commands still write
StateBench result directories if executed, so use generated commands with
--list-tasks when only checking suite selection.
Target-suite validation:
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--suite-target-report
Readiness report:
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--readiness-report
Artifact preflight template:
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--preflight-template \
--artifact-id qwen3-6-35b-a3b-hf
Write and validate preflight records:
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--manifest statebench/suites/finance-model-runtime-candidates-v0.json \
--preflight-template \
--write-preflight-template \
--artifact-id qwen3-6-35b-a3b-hf
scripts/.venv-tts/bin/python3 -m statebench.runners.model_candidate_manifest \
--validate-preflight-record statebench/preflights/qwen3-6-35b-a3b-hf-preflight.json
The validator enforces
statebench/schemas/preflight_record.schema.json
before checking whether all required preflight items have evidence.
Scored finance-replay-v0 suite runs and finance-retrieval-sql-harness-v0
backend calls require --preflight-record; dry-runs, --list-tasks, fixture
smoke checks, prompt export, Postgres preflight, and scoring an existing answer
artifact remain available without approval.
- Run a frontier API control against all 118
finance-replay-v0tasks. - Run
Qwen/Qwen3.6-35B-A3Bthrough an OpenAI-compatible vLLM or SGLang endpoint, with exact HF revision, 262K effective context unless extended context is deliberately enabled, and a thinking-mode note. - Run
google/gemma-4-31Bor the best Gemma 4 artifact that fits available hardware; record whether the run is dense, MoE, API, GGUF, or MLX. - Run one Liquid task-specialist artifact on the subset it should plausibly improve: source-pack QA, transcript cleanup, extract/PII, RAG, tool routing, or reranking. Do not force it into general finance reasoning.
- Run one finance specialist only after its license and artifact are verified.
rLLM-FinQA-4B, ODA-Fin-style Qwen3 finance post-training, andTheFinAI/Fin-o1-8Bare higher signal than old BloombergGPT/PIXIU references because they are closer to local StateBench experiments.khazarai/Fino1-4Bis now the cheaper Qwen3-4B FinQA-style control in the same specialist lane, but it remains blocked on license-conflict and runtime preflight checks. 5a. Do not leave new static-control suites without candidate artifacts. The manifest now targetsfinance-treasury-payments-liquidity-v0,finance-agentic-factor-discovery-v0, andfinance-portfolio-risk-construction-v0with frontier/API ceiling control, Qwen3.6 open-base comparison, Gemma 4 where the task is factor/portfolio reasoning, and Liquid Financial Analyst only as a treasury/payments tiny-model cost-floor control. These are planning rows, not model scores. 5b. Treat Trading-R1 / Alpha-R1 / Trade-R1 as a separate paper-first trading/factor RL lane. They are highly relevant to systematic quant research because they target structured trading theses, factor screening, noisy market rewards, and reward hacking, but they should not enter a scored runtime queue until code, weights, license, exact revision, backend, transaction-cost assumptions, point-in-time data, and reproducibility evidence are recorded. - Use
finance-sentiment-routing-v0before using large chat models for headline or earnings-snippet classification. The 15-task held-out fixture and static gold-control scorecard now exist. FinSenti-Qwen3.5-9B is the current Qwen-family candidate to compare against FinBERT/ModernFinBERT and tiny Qwen3-0.6B sentiment controls such asYorkFr/financial-sentiment-qwen3-v2, but it is not runnable in this repo until the Qwen3.5 image-text-to-text adapter, FinSenti tag parser, or llama.cpp/GGUF backend is proven. 6a. Add an earnings-call evasion sublane. Eva-4B-V2 is a Qwen3-4B-derived specialist for analyst-question / management-answer directness and evasiveness. It belongs in transcript/disclosure-quality triage and earnings-call feature extraction, not in generic sentiment, alpha, investment advice, or portfolio construction. Any score must record label-source bias, LLM-as-judge risk, intermediate-class subjectivity, data-vintage drift, and human-review requirements. - Build the accounting/compliance guard lane separately from broad finance reasoning. CPA-Qwen3 belongs on accounting/tax/audit boundary tasks, while FinGuard belongs on policy-grounded verifier tasks after jurisdiction and policy corpus are defined.
- Build the multimodal finance lane separately from the text patch runner.
AITRADER/Amsi-fin-o1.5-mxfp8-MLXis useful for chart/OCR/screenshot source tasks only after anmlx-vlmsmoke test proves it runs locally. - Build the retrieval lane separately from the patch runner. Start with Qwen3 Embedding plus Qwen3 Reranker against FinRetrieval-style and internal source-localization tasks, then compare finance-specific/gated encoders such as FinE5 if licensing permits.
- Build the SQL lane separately from the patch runner. Treat Snowflake Arctic-Text2SQL-R2 as the current Snowflake-specialist recipe to test, but do not score vendor-post claims as internal performance.
- Only after base/API/local runs identify repeatable failure modes, compare Unsloth, LLaMA-Factory, AWS native, and MLX fine-tuning paths.
Required Manifest Metadata
Every run must fill these fields in statebench/results/<suite_run_id>/manifest.json:
modelweights_revisionartifact_idadapter_idmerge_stateartifact_formatartifact_precisionbackend_idbackend_hardwarebackend_constraintscontext_length
For backend constraints, record negative knowledge explicitly, for example:
neuron compatibility unproven until operator and compilation preflight passesmlx conversion not testedgguf quantization is community artifact, not upstream releaseapi-only proprietary routelicense review required before commercial use
Why This Matters For The Finance Goal
The repo already has broad finance/quant coverage. The remaining risk is selection drift: choosing models because they are fashionable, locally convenient, or vendor-promoted instead of because they pass the finance tasks we actually need.
This registry prevents three common mistakes:
- Treating Qwen, Gemma, or Liquid family names as scores.
- Letting SQL/retrieval/source-discovery performance leak into the general patch-runner benchmark.
- Assuming deployment portability across CUDA, MLX, llama.cpp, AMD, and Neuron without explicit artifact evidence.
Current Interpretation
As of May 29, 2026:
Qwen/Qwen3.6-35B-A3Bis the default current open Qwen-family base candidate for StateBench finance replay. It is current enough to avoid the stale Qwen2.5 trap and has explicit long-context and agentic-coding evidence.- Gemma 4 is the right open-family counterweight to Qwen, especially for multimodal and long-context document tasks.
- LiquidAI is most interesting as a family of small task specialists and portable runtime artifacts, not as a finance alpha/reasoning model.
- Arctic-Text2SQL-R2 is the most important new SQL signal because it argues that compact, domain/dialect-trained models can beat frontier systems on enterprise SQL. That belongs in a Snowflake/Postgres harness, not the general finance replay patch runner.
Dev9124/qwen3-finance-modelis a current Qwen3 finance-instruction watchlist artifact with Apache-2.0 model-card licensing and vLLM/SGLang runtime snippets. It is useful as a broad fine-tune control, but model-card evidence alone is weaker than ODA-Fin/rLLM-style paper-backed specialists.TheFinAI/Fin-o1-8Bis now tracked as a Qwen3-8B finance-reasoning candidate. The current HF card supersedes the stale local Llama-lineage note; score it as a concrete Qwen3 artifact only after exact revision and runtime proof.khazarai/Fino1-4Bis now tracked as a separate Qwen3-4B finance-reasoning watchlist candidate, not a smaller alias forTheFinAI/Fin-o1-8B. The captured HF revision is15c0d969d0bd2aa6515e0e315b5aecc063c2d7ee, architecture isQwen3ForCausalLM, base metadata isunsloth/Qwen3-4B, and the model has 4,022,468,096 BF16 safetensors parameters. Its dataset points toTheFinAI/Fino1_Reasoning_Path_FinQAat revision0316559663c46e50d90e75f6e14a6d8f679cdba8undercc-by-4.0. Block production/commercial assumptions until the model license conflict is resolved: HF metadata says Apache-2.0 while the README body says MIT. Treat it as a cheap FinQA-style numerical/table-QA control only; no GGUF, MLX, ROCm, SageMaker, Bedrock, or Neuron portability is proven.- CPA-Qwen3 and FinGuard widen the domain-specialist lane from investment
reasoning into accounting/compliance and policy-grounded guardrails. Their
scores should live on accounting/compliance fixtures, not generic replay. The
first
finance-accounting-compliance-v0seed fixture now exists with 10 guardrail tasks and a static10/10answer-artifact control; model/backend scores still require preflight approval. AITRADER/Amsi-fin-o1.5-mxfp8-MLXadds a current MLX finance VLM candidate for charts, OCR, screenshots, and multimodal source triage. It is not proof that MLX should replace other runtimes, and it is not a text-only benchmark default.YorkFr/financial-sentiment-qwen3-v2is a tiny MIT-licensed Qwen3-0.6B sentiment artifact. It belongs in a held-out sentiment/routing lane, not the general finance replay patch runner. The fixture and Transformers-local adapter now exist. The first local run is negative evidence:0/15accepted, despite partial label signal, because route/rationale controls failed.- Regional and narrow finance SLMs are now tracked as a separate watchlist:
iravikr/qwen3-0.6b-finance-indiafor India budget/policy QA,KBTG-Labs/THaLLE-0.2-ThaiLLM-8B-fafor Thai/English finance-language coverage, andAyansk11/qwen3-4b-financial-sentiment-grpoas a smaller Qwen3 sentiment/GGUF candidate.sweatSmile/Qwen3-4B-Instruct-FinanceQAis a useful cautionary control because its Hugging Face metadata reports a Qwen2.5 base despite the repository name. - The May 31 HF watchlist pass adds a more concrete local runtime candidate:
mradermacher/Qwen3-8B-finance-V1.0.1-Merged-V2-GGUF, revision29d21ffe58d0e9c7f86136fe7f9eaab3eb0d1332, as a llama.cpp/GGUF finance reasoning control. It should be scored as its exact quant file, not as the upstream RinKana merge. The same pass addsPeterAM4/Qwen3-Embedding-0.6B-GGUF, revisionbb661aeeeafa4ff7303b5f5e1e80255efd6a513c, as a tiny local embedding candidate with finance-related imatrix/data tags. - The June 5 HF watchlist follow-up adds
shm2290/finance-embeddings-gemma-300m-v2as a finance-specific EmbeddingGemma 300M control. It is useful precisely because it is small and retrieval-only: test it on source localization, wiki/static-index retrieval, SEC/investor-document paragraph retrieval, and podcast/source-card metadata search. Do not use it as a finance reasoner, source-discovery agent, accounting model, or alpha signal. Its license inherits fromgoogle/embeddinggemma-300m, so commercial and redistribution use need base license review before promotion. - The June 6 HF watchlist follow-up adds
souflex56/qanchor-reranker-qwen3-0.6b-mergedas a Chinese financial-document reranker candidate. It is useful because it targets A-share annual reports and related filings with a current Qwen3 sequence-classification reranker, but it is not a generator or broad finance model. Score it only after a pinned Chinese financial-document retrieval fixture, first-stage BM25/embedding/RRF baseline, preservedqwen3_templateformatting, same-document scope check, and reranker trace schema are in place. - The June 6 regional-sentiment follow-up adds
p988744/eland-sentiment-zh-vllmas a Qwen3-4B Traditional Chinese/Taiwan financial-sentiment routing candidate. It is useful for Taiwan stock-market headlines, entity sentiment, and opinion sentiment, but it is not a U.S. SEC, accounting, quant-research, investment-advice, portfolio, trading, or alpha model. Score only after exact HF revisions are captured, the required system prompt is preserved, a held-out Traditional Chinese/Taiwan fixture exists, and vLLM/PEFT/GGUF artifacts are separated. Do not use Ollama. - The June 6 structured-extraction follow-up adds
lmxxf/financial-report-lora-qwen3-14bas a Chinese financial research-report metric-extraction adapter onQwen/Qwen3-14B. It is useful because it targets schema-checked extraction of deeply analyzed metrics from Chinese report paragraphs, a better fit for source cards and factor-discovery bootstrapping than generic sentiment. It is blocked on license visibility, base revision, PEFT adapter-load smoke, held-out Chinese report metric spans, JSON-schema validation, and output-hygiene checks. Do not treat it as a broad finance reasoner, accounting authority, alpha model, or runtime portability proof. Mikkkkoooo/qwen35-4b-private-analyst-full-corpusandDMindAI/DMind-3-miniare recorded as watchlist controls only. The former is a private-corpus analyst-style fine-tune with licenseother; the latter is a Web3/DeFi finance model, not a U.S. asset-management or SEC-filing model.- The May 31 follow-up pass adds three narrow candidates. Mihenk is a
Qwen3.6-35B-A3B Turkish/BIST finance model now covered by the
finance-regional-language-v0scope-control fixture. Distil/LFM2.5 voice assistant is a 354M banking tool-calling candidate for voice/tool-routing fixtures, not quant research. TokenFactory Gemma SEC extraction v3 is a GGUF/llama.cpp structured-extraction candidate with unverified card metrics and Gemma-license constraints. - The June 5 follow-up records a separate community GGUF conversion of Mihenk. This improves the local llama.cpp runtime lane but does not change the upstream HF/BF16 evidence. The GGUF artifact is blocked until an exact file, prompt-template compatibility, and llama.cpp smoke test are recorded.
- The June 5 Gemma follow-up records
Sengil/turkish-gemma-9b-finance-sftas a regional Turkish finance-language Gemma2 adapter. It is useful as a family-control against the Qwen/Mihenk Turkish lane, but it remains blocked on base-license review, adapter-load smoke, and visible reasoning/output-hygiene checks. - The second May 31 follow-up pass adds three more executable local controls.
Myyyyyyyyyyyyyy/qwen3-14b-bookkeeper-gguffills a bookkeeping/accounting workflow gap with Qwen3-14B GGUF and explicit work-in-progress caveats.mad-lab-ai/qwen3-1.7b-sentiment-ggufis a tiny local sentiment-cost-floor candidate for label/routing fixtures, with leakage and alpha-overclaim guardrails.vividdream/Qwen-Open-Finance-R-8B-IQ4_NL-GGUFgives a multilingual finance/economics llama.cpp control that must be scored separately from upstream DragonLLM weights and other quantizations. neoyipeng/ModernFinBERT-baseandProsusAI/finbertestablish label-only encoder baselines on the same fixture. ModernFinBERT currently leads the label slice (0.733vs.0.667for ProsusAI and YorkFr), but both encoders fail full routing acceptance because they do not emit route/rationale policy.- AssetOpsBench QLoRA tool-knowledge evidence strengthens the fine-tuning gate: small models may internalize stable tool catalogs, but only after StateBench shows repeated tool/planning failures and after catastrophic forgetting is measured.
- Qwen3.7-Max is worth tracking as a proprietary long-horizon agent reference, but it is not an open/local candidate and should not displace Qwen3.6 for the local artifact lane without verified open weights.
- Relational Probing adds a distinct market-graph lane: Qwen3 0.6B/1.7B/4B SLMs are used as upstream hidden-state encoders whose relation head induces financial-entity graphs for downstream stock-trend prediction. This is closer to quant modeling than finance chat, but it is paper-first evidence until code/data/checkpoints are verified and leakage-safe point-in-time fixtures exist.