Yi Wang
CokeWang
AI & ML interests
Agent, Time-series, LLM, Multimodal LLM
Recent Activity
posted an update about 8 hours ago
LoopArena: Which Models Make Good Runtime Controllers for Coding Agents?
Long-running coding agents are often guided by another model that reviews progress, chooses the next assignment, requests verification, and decides when to stop. LoopArena evaluates that model as the Controller while keeping the coding Worker and execution setup fixed.
The benchmark covers next-step decisions, repeated control over task slices, and complete software tasks. We have released the benchmark data, evaluation code, and v0.1.0 results.
https://huggingface.co/papers/2608.28281
https://github.com/AMAP-ML/LoopArena upvoted a paper 4 days ago
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution submitted a paper 4 days ago
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering