arxiv:2610.01509
🔄 In a Training Loop
Changdae Oh
changdae
AI & ML interests
Generalization; Distribution Shift; Uncertainty Quantification; Reward Modeling; Post-training
Recent Activity
upvoted a paper 2 days ago
MIMESIS: Learning User Simulators as Training Environments for Interactive Agents upvoted a paper 3 days ago
On-Policy Distillation with Negative-Policy Rollouts upvoted a paper 4 days ago
DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling