machinelearning.apple.com·9d ago
From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers
Submitted by @

Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holistic scalar objective. We introduce a rubric-based reward framework that generates query-specific rubrics grounded in retrieved evidence and decomposed into multiple quality dimensions, providing fine-grained supervision during post-training. Averaged across three evaluation axes
Apple Machine Learning Research
Published 9d ago ago
Original reporting by Apple Machine Learning Research. GridIndex is an aggregation and intelligence layer — full credit to the original publisher.
This page summarizes and tracks coverage of this developing story.
Read full story at Apple Machine Learning Research 0 views 0 upvotes 0 comments 0 shares
Discussion · 0
Sign in to join the discussion.