ProactiveInquirer-Qwen3-8B

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM   2Weizmann Institute of Science

Project page Paper Code License

▶ The paper's example, step by step (22 seconds): the questioner trained with Q&D finds the account, the order with the boots and the size-8 boots, and the task is completed.

This is the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents. It is a LoRA adapter on Qwen3-8B, trained with Q&D (questioner and drafter).

A tool-using agent usually does what it is asked, yet a task often needs information the user never mentions. This model decides, one step at a time, which question to send to the agent's retriever next, or that it is time to stop. It goes after two kinds of unstated need:

  • Horizontal proactivity: a need the current state already names. A customer gives a name and a ZIP code, so the agent looks up the account.
  • Vertical proactivity: a need that only newly found evidence names. The account lists the order, the order names the product, and the product lists the size-8 variant.

One customer request, the same model prompted and trained

Results

These numbers are from the paper, on held-out test splits. The MuSiQue reading compares policies after the same number of questions, so asking more cannot pass for asking better. The τ²-bench readings are the benchmark's own task success under the same per-dialogue caps.

Setting Measure Qwen3-8B, prompted This model
MuSiQue, equal retrieval spend Required evidence recovered 78% 90%
τ²-bench retail (base prompt), no further training Task success 13% 34%
τ²-bench retail (stop prompt), no further training Task success 12% 32%
  • At equal retrieval spend it improves both forms of proactivity over the same model, prompted, on held-out splits of three multi-hop QA benchmarks, and outperforms GPT-OSS-120B, a prompted model 15× larger in the same role, on two of the three.
  • The gain comes from what it asks, not from asking more or longer questions: it holds against a question-volume control and a length control.
  • Placed in a customer-service agent with a simulated customer, with no further training, it raises retail task success from 13% to 34%, and it asks less and finds more: fewer questions, more of which reach the records the task needs. Against GPT-OSS-120B in retail, it completes more tasks with fewer follow-up turns from the customer.

Required-evidence coverage against retrieval calls

How to use it

The questioner reads one prompt, the template it was trained on, and replies with one JSON action. The two template files are in prompts/.

import re

import torch
from huggingface_hub import hf_hub_download
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "dolev31/ProactiveInquirer-Qwen3-8B"
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3-8B")
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B", dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, REPO)  # seed 1; add subfolder="seed2" for the second seed

template = open(hf_hub_download(REPO, "prompts/inquirer_prompted.txt"), encoding="utf-8").read()
placebo = open(hf_hub_download(REPO, "prompts/fragment_user_channel_placebo.txt"), encoding="utf-8").read()


def next_action(**state):
    fields = dict(state, user_channel=placebo.strip())
    prompt = re.sub(r"\{\{(\w+)\}\}", lambda m: str(fields[m.group(1)]), template)
    ids = tok.apply_chat_template(
        [{"role": "user", "content": prompt}],
        add_generation_prompt=True,
        enable_thinking=False,
        return_tensors="pt",
        return_dict=True,
    ).to(model.device)
    out = model.generate(**ids, max_new_tokens=200, do_sample=False)
    return tok.decode(out[0, ids["input_ids"].shape[1] :], skip_special_tokens=True)


question = "Who was the spouse of the director of the film The Great Flamarion?"
instructions = (
    "Answer the question using a closed pool of 20 paragraphs. You may issue retrieval "
    "queries against that pool before answering; several paragraphs are distractors, and "
    "the answer usually requires composing facts from more than one of them."
)
print(next_action(
    question=question, instructions=instructions, evidence="(nothing retrieved yet)",
    draft="(no draft yet)", history="(nothing asked yet)",
))
evidence = (
    "[3f2a9c1b7d4e] The Great Flamarion\n"
    "The Great Flamarion is a 1945 American film noir directed by Anthony Mann and starring "
    "Erich von Stroheim, Mary Beth Hughes and Dan Duryea."
)
print(next_action(
    question=question, instructions=instructions, evidence=evidence,
    draft="The film was directed by Anthony Mann; his spouse is not yet known.",
    history="Q1: Who directed the film The Great Flamarion?\n"
            "A1: The Great Flamarion (1945) was directed by Anthony Mann.",
))

Output, with greedy decoding (run on one A100 with transformers 5.17, peft 0.20 and torch 2.14):

{"action": "ASK", "question": "Who directed the film The Great Flamarion?", "rationale": "Identify the director to later find their spouse"}
{"action": "ASK", "question": "Who was the spouse of film director Anthony Mann?", "rationale": "Need the spouse of the director to answer the task"}

The first question is horizontal: the task names the film, so the questioner goes after its director. The second is vertical: it uses a value only the retrieved evidence named, Anthony Mann. Compare the action field case-insensitively, since the model may write ASK or ask.

Inside an agent, feed each answer back: retrieved paragraphs go into evidence as [uid] title plus text, questions and answers into history as Q1:/A1: lines, and the drafter's text into draft. The ProactiveInquirer library runs the whole loop (questioner, retriever, drafter, answerer) and the paper's evaluation. The project page summarizes the paper and its results.

To serve it with vLLM, download the adapter and pass it as a LoRA module (rank 32):

huggingface-cli download dolev31/ProactiveInquirer-Qwen3-8B --local-dir proactive-inquirer
vllm serve Qwen/Qwen3-8B --enable-lora --max-lora-rank 32 \
    --lora-modules proactive-inquirer=./proactive-inquirer

Other formats

Training

  • Method. Q&D trains the questioner from the consequences of its own questions. A run is forked at one state and continued after several candidate questions (and after stopping). The candidate whose continuation retrieves more of the required evidence is preferred, and asking is preferred over stopping while evidence is still missing. The drafter is frozen, so every change in what the agent holds is caused by a question. No reward model or model judge is involved.
  • Stages. Three. First imitation: where the required evidence was already in hand the target is to stop, and otherwise the best sampled question, if its consequence score clears a fixed floor. Then direct preference optimization on question pairs, and last on question pairs and stop contrasts together, which rank asking above stopping at unfinished states. This adapter is the last stage.
  • Data. Training splits of MuSiQue, StrategyQA and 2WikiMultiHopQA. The final stage uses 31,473 preference pairs over 19,124 states (11,305 question-vs-question, 1,249 question-vs-stop and 18,919 synthetic question-vs-stop pairs). Nothing from τ²-bench is used in training.
  • Hyperparameters. LoRA r = 32, α = 64, dropout 0.05 on all attention and MLP projections. DPO with the sigmoid loss, β = 0.1, learning rate 5e-6, one epoch, gradient accumulation 16, bf16, sequences up to 5,120 tokens, Qwen3 chat template with thinking disabled.
  • Seeds. The paper reports two training seeds: seed 1 is at the root of this repository, seed 2 in seed2/.

Limitations

  • It has learned what to ask more readily than when to stop.
  • The extra evidence it finds does not yet translate into better final answers.
  • User-facing results come from a simulated customer, not from real people.
  • It is a component inside an agent, meant to be called with the template above. It is not a chat assistant, and it was trained and evaluated in English.

Citation

@article{levy2026asking,
  title   = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
  author  = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
  journal = {arXiv preprint arXiv:2609.37236},
  url     = {https://arxiv.org/abs/2609.37236},
  year    = {2026}
}

License

Apache-2.0, like the base model Qwen3-8B.

Downloads last month
57
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dolev31/ProactiveInquirer-Qwen3-8B

Finetuned
Qwen/Qwen3-8B
Adapter
(2238)
this model

Datasets used to train dolev31/ProactiveInquirer-Qwen3-8B

Collection including dolev31/ProactiveInquirer-Qwen3-8B

Paper for dolev31/ProactiveInquirer-Qwen3-8B