File size: 2,434 Bytes
4258b26
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5f62acb
4258b26
5f62acb
 
4258b26
 
 
 
 
 
 
 
 
 
5f62acb
4258b26
 
5f62acb
4258b26
 
5f62acb
4258b26
5f62acb
4258b26
 
5f62acb
4258b26
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5f62acb
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
base_model: google/gemma-3-4b-it
base_model_relation: finetune
library_name: peft
license: mit
tags:
  - lora
  - peft
  - transformers
  - audio-question-answering
  - correctness-assessment
  - orca
language:
  - en
pipeline_tag: text-classification
---

# ORCA — Gemma-3-4B-IT (Multinomial, seed 99)

ORCA (**O**pen-ended **R**esponse **C**orrectness **A**ssessment) scores the correctness of open-ended audio QA responses. Given a question, reference answer, candidate answer, and an LLM-generated rationale, it outputs a correctness score in [0, 1] and an uncertainty estimate.

**Paper:** [ORCA: Open-ended Response Correctness Assessment for Audio Question Answering](https://arxiv.org/abs/2512.09066) — accepted to *TACL 2026*  
**Code & usage:** [github.com/BUTSpeechFIT/ORCA](https://github.com/BUTSpeechFIT/ORCA)  
**Training data:** [BUT-FIT/orca-audio-qa-annotations](https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations)

## Model details

| Property | Value |
|---|---|
| Base model | `google/gemma-3-4b-it` |
| LoRA rank / alpha | 128 / 128 |
| Loss function | Multinomial log-likelihood (5-class Likert) |
| Training seed | 99 |
| Training curriculum | Stage 1 (synthetic) → Stage 2 (LLM-judge) → Stage 3 (human) |
| Precision | bfloat16 |

## Quick start

```bash
pip install git+https://github.com/BUTSpeechFIT/ORCA.git
hf download BUT-FIT/orca-gemma-3-4b-it-multinomial --local-dir orca-gemma-4b
orca-infer --model_path orca-gemma-4b/model --data_jsonl your_data.jsonl --output_dir results/
```

See the [repository](https://github.com/BUTSpeechFIT/ORCA) for full usage, evaluation scripts, and the `download_and_infer.py` convenience script.

## Citation

```bibtex
@article{sedlacek-etal-2026-orca,
  title={ORCA: Open-ended Response Correctness Assessment for Audio Question Answering},
  author={Sedl\'{a}\v{c}ek, \v{S}imon and Barahona, Sara and Bola\~{n}os, Cecilia and
          Herrera-Alarc\'{o}n, Laura and Udupa, Sathvik and L\'{o}pez, Fernando and
          Ferner, Allison and Lozano-Diez, Alicia and Yusuf, Bolaji and Kesiraju, Santosh and
          Duraiswami, Ramani and \v{C}ernock\'{y}, Jan},
  howpublished={Accepted to Transactions of the Association for Computational Linguistics},
  year={2026},
  url={https://arxiv.org/abs/2512.09066}
}
```

## License

MIT License. See the [repository LICENSE](https://github.com/BUTSpeechFIT/ORCA/blob/main/LICENSE) for details.