Instructions to use dolev31/ProactiveInquirer-Qwen3-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use dolev31/ProactiveInquirer-Qwen3-8B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B") model = PeftModel.from_pretrained(base_model, "dolev31/ProactiveInquirer-Qwen3-8B") - Notebooks
- Google Colab
- Kaggle
ProactiveInquirer-Qwen3-8B
Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM 2Weizmann Institute of Science
▶ The paper's example, step by step (22 seconds): the questioner trained with Q&D finds the account, the order with the boots and the size-8 boots, and the task is completed.
This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained with Q&D (questioner and drafter).
A tool-using agent usually does what it is asked, yet a task often needs information the user never mentions. This model decides, one step at a time, which question to send to the agent's retriever next, or that it is time to stop. It goes after two kinds of unstated need:
- Horizontal proactivity: a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.
- Vertical proactivity: a need that only newly found evidence names. The account lists the order, the order names the product, and the product lists the size-8 variant.
Results
These numbers are from the paper, on held-out test splits. The MuSiQue reading compares policies after the same number of questions, so asking more cannot pass for asking better. The τ²-bench readings are the benchmark's own task success under the same per-dialogue caps.
| Setting | Measure | Qwen3-8B, prompted | This model |
|---|---|---|---|
| MuSiQue, equal retrieval spend | Required evidence recovered | 78% | 90% |
| τ²-bench retail (base prompt), no further training | Task success | 13% | 34% |
| τ²-bench retail (stop prompt), no further training | Task success | 12% | 32% |
- At equal retrieval spend it improves both forms of proactivity over the same model, prompted, on held-out splits of three multi-hop QA benchmarks, and outperforms GPT-OSS-120B, a prompted model 15× larger in the same role, on two of the three.
- The gain comes from what it asks, not from asking more or longer questions: it holds against a question-volume control and a length control.
- Placed in a customer-service agent with a simulated customer, with no further training, it raises retail task success from 13% to 34%, and it asks less and finds more: fewer questions, more of which reach the records the task needs. Against GPT-OSS-120B in retail, it completes more tasks with fewer follow-up turns from the customer.
How to use it
The questioner reads one prompt, the template it was trained on, and replies with one JSON action. The
two template files are in prompts/.
import re
import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "dolev31/ProactiveInquirer-Qwen3-8B"
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, REPO) # seed 1; add subfolder="seed2" for the second seed
template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()
def next_action(**state):
fields = dict(state, user_channel=placebo.strip())
prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
ids = tok.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt",
return_dict=True,
).to(model.device)
out = model.generate(**ids, max_new_tokens=200, do_sample=False)
return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)
question = "Who was the spouse of the director of the film The Great Flamarion?"
instructions = (
"Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
"queries against that pool before answering; several paragraphs are distractors, and "
"the answer usually requires composing facts from more than one of them."
)
print(next_action(
question=question, instructions=instructions, evidence="(nothing retrieved yet)",
draft="(no draft yet)", history="(nothing asked yet)",
))
evidence = (
"[3f2a9c1b7d4e] The Great Flamarion\n"
"The Great Flamarion is a 1945 American film noir directed by Anthony Mann and starring "
"Erich von Stroheim, Mary Beth Hughes and Dan Duryea."
)
print(next_action(
question=question, instructions=instructions, evidence=evidence,
draft="The film was directed by Anthony Mann; his spouse is not yet known.",
history="Q1: Who directed the film The Great Flamarion?\n"
"A1: The Great Flamarion (1945) was directed by Anthony Mann.",
))
Output, with greedy decoding (run on one A100 with transformers 5.17, peft 0.20 and torch 2.14):
{"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
{"action": "ASK", "question": "Who was the spouse of film director Anthony Mann?", "rationale": "Need the spouse of the director to answer the task"}
The first question is horizontal: the task names the film, so the questioner goes after its director.
The second is vertical: it uses a value only the retrieved evidence named, Anthony Mann. Compare the
action field case-insensitively, since the model may write ASK or ask.
Inside an agent, feed each answer back: retrieved paragraphs go into evidence as [uid] title plus
text, questions and answers into history as Q1:/A1: lines, and the drafter's text into draft.
The ProactiveInquirer library runs the whole loop
(questioner, retriever, drafter, answerer) and the paper's evaluation. The
project page summarizes the paper and its results.
To serve it with vLLM, download the adapter and pass it as a LoRA module (rank 32):
huggingface-cli download dolev31/ProactiveInquirer-Qwen3-8B --local-dir proactive-inquirer
vllm serve Qwen/Qwen3-8B --enable-lora --max-lora-rank 32 \
--lora-modules proactive-inquirer=./proactive-inquirer
Other formats
- Merged weights, which load without PEFT and serve like any Qwen3-8B: dolev31/ProactiveInquirer-Qwen3-8B-Merged.
- GGUF for llama.cpp, Ollama and LM Studio (Q4_K_M, Q5_K_M, Q8_0):
dolev31/ProactiveInquirer-Qwen3-8B-GGUF.
To run it locally:
ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --think=false.
Training
- Method. Q&D trains the questioner from the consequences of its own questions. A run is forked at one state and continued after several candidate questions (and after stopping). The candidate whose continuation retrieves more of the required evidence is preferred, and asking is preferred over stopping while evidence is still missing. The drafter is frozen, so every change in what the agent holds is caused by a question. No reward model or model judge is involved.
- Stages. Three. First imitation: where the required evidence was already in hand the target is to stop, and otherwise the best sampled question, if its consequence score clears a fixed floor. Then direct preference optimization on question pairs, and last on question pairs and stop contrasts together, which rank asking above stopping at unfinished states. This adapter is the last stage.
- Data. Training splits of MuSiQue, StrategyQA and 2WikiMultiHopQA. The final stage uses 31,473 preference pairs over 19,124 states (11,305 question-vs-question, 1,249 question-vs-stop and 18,919 synthetic question-vs-stop pairs). Nothing from τ²-bench is used in training.
- Hyperparameters. LoRA r = 32, α = 64, dropout 0.05 on all attention and MLP projections. DPO with the sigmoid loss, β = 0.1, learning rate 5e-6, one epoch, gradient accumulation 16, bf16, sequences up to 5,120 tokens, Qwen3 chat template with thinking disabled.
- Seeds. The paper reports two training seeds: seed 1 is at the root of this repository, seed 2 in
seed2/.
Limitations
- It has learned what to ask more readily than when to stop.
- The extra evidence it finds does not yet translate into better final answers.
- User-facing results come from a simulated customer, not from real people.
- It is a component inside an agent, meant to be called with the template above. It is not a chat assistant, and it was trained and evaluated in English.
Citation
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint arXiv:2609.37236},
url = {https://arxiv.org/abs/2609.37236},
year = {2026}
}
License
Apache-2.0, like the base model Qwen3-8B.
- Downloads last month
- 57

