Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Wang
WayenVan
2
1
4
Follow
HintonZhang50's profile picture
winniemangeni's profile picture
2 followers
ยท
3 following
AI & ML interests
None yet
Recent Activity
updated
a dataset
about 6 hours ago
WayenVan/rnb-radar-v2.0
published
a dataset
about 6 hours ago
WayenVan/rnb-radar-v2.0
reacted
to
Kseniase
's
post
with ๐ฅ
11 months ago
10 Latest Preference Optimization Techniques Models need feedback on what makes outputs โgoodโ or โbad.โ Policy optimization (PO) turns preferences and rewards into actual training signals. This field is evolving quickly, moving far beyond classics like PPO and GRPO. So here is our overview of 10 newest PO methods: 1. Pref-GRPO โ https://huggingface.co/papers/2508.20751 Stabilizes text-to-image reinforcement learning (RL) with pairwise preference rewards and a unified UNIGENBENCH benchmark 2. PVPO (Policy with Value Preference Optimization) โ https://huggingface.co/papers/2508.21104 This critic-free RL method uses a pre-trained model as a reference anchor to reduce bias and guide learning, selecting high-value examples through data pre-sampling 3. DCPO (Dynamic Clipping Policy Optimization) โ https://huggingface.co/papers/2509.02333 Uses dynamic clipping, which adjusts probability limits per token for better token exploration, and smooth reward standardization to balance rewards over training steps and prevent wasted updates 4. ARPO (Agentic Reinforced Policy Optimization) โ https://huggingface.co/papers/2507.19849 Optimizes multi-turn LLM agents that use external tools. It uses an entropy-based adaptive rollout to explore post-tool use and an advantage attribution method to better assign credit across steps, leading to more efficient tool use with fewer resources 5. GRPO-RoC (Group Relative Policy Optimization with Resampling-on-Correct) โ https://huggingface.co/papers/2508.20722 Oversamples rollouts, then resamples them to keep diverse mistakes and only the highest-quality correct answers. It reduces noises and ends up with stronger reasoning in a code environment Read further below โฌ๏ธ If you like this, also subscribe to the Turing post: https://www.turingpost.com/subscribe
View all activity
Organizations
WayenVan
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
WayenVan/PHOENIX-Weather14T
12 months ago
[bot] Conversion to Parquet
#1 opened about 1 year ago by
parquet-converter
New activity in
google/gemma-3-1b-it
about 1 year ago
Why are vocab_size and tokenizer different length?
๐
5
7
#17 opened over 1 year ago by
choco9966