FinEdgar Gemma 4 E2B LoRA

This repository contains a LoRA adapter fine-tuned for question answering over SEC filings.

The adapter is intended to be used inside a retrieval-augmented financial QA system. It was trained to answer with filing-grounded context and to work alongside deterministic XBRL tooling for structured financial facts.

Base Model

  • Base model: google/gemma-4-e2b-it
  • Adapter type: LoRA
  • PEFT version: 0.18.1
  • Task type: causal language modeling

Users must have access to the base model and comply with the base model license and terms.

Intended Use

This adapter is intended for:

  • SEC filing question answering
  • filing-grounded financial summaries
  • structured financial QA when paired with XBRL facts
  • local RAG pipelines over 10-K, 10-Q, and 8-K filings

It is not intended to be used as a standalone source of financial truth. Numeric answers should be checked against SEC XBRL facts or the original filing.

Out-of-Scope Use

Do not use this model as financial advice, investment advice, accounting advice, legal advice, or a substitute for reviewing original SEC filings.

The model can be wrong when retrieval context is missing, incomplete, stale, or ambiguous.

Training Data

Training used a mixture of:

  • financial QA examples
  • financial sentiment examples
  • SEC/XBRL-derived synthetic QA
  • filing-diff summary examples

FinanceBench was used for evaluation and was not used as direct training supervision.

Evaluation

Latest tracked FinanceBench result from the project evaluation run:

Metric Result
Overall accuracy 40.7%
XBRL route accuracy 40.0%
RAG route accuracy 41.1%
Recall@5 30.0%
Average latency 2,144 ms

Evaluation file: eval_latest_model_20260414.json

These results are system-level results from the full FinEdgar pipeline, not adapter-only generation in isolation.

Limitations

  • The adapter depends on good retrieval context.
  • It can produce incomplete narrative answers when evidence is spread across multiple filings or sections.
  • It may format or reason about financial metrics incorrectly without deterministic XBRL support.
  • It does not know current market prices or events unless those are provided in context.
  • It should be used with citations and source checking.

Loading

Example PEFT loading pattern:

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = "google/gemma-4-e2b-it"
adapter = "YOUR_USERNAME/finedgar-gemma-4-e2b-lora"

tokenizer = AutoTokenizer.from_pretrained(adapter)
model = AutoModelForCausalLM.from_pretrained(base_model)
model = PeftModel.from_pretrained(model, adapter)

Depending on the installed Transformers version and Gemma 4 support, a model-specific class or processor may be required instead of AutoModelForCausalLM.

Suggested Inference Pattern

Use the model with retrieved filing context:

Question: <user question>

Context:
<retrieved SEC filing excerpts with citations>

Answer using only the provided context. If the answer is not present, say so.

For structured metrics, prefer XBRL lookup and use generation only for explanation or formatting.

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support