Reinforcement Learning
7 tracked items
Fortune·9d ago
Anthropic pauses some AI training following rogue agent hacks. Here’s how its compares to OpenAI’s.
arXiv·9d ago
Cognitively-Grounded On-Device Runtime Learning for Ground Robots in Unknown Physical Environments
arXiv·9d ago
AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning
arXiv·9d ago
Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization
arXiv·9d ago
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
arXiv·9d ago
StreamScout: Learning When to Look Deeper for Streaming Video Understanding
arXiv·9d ago
Teaching Robot Policies to Humans Using Erroneous Examples