Benchmarks
22 tracked items
arXiv·9d ago
A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving
arXiv·9d ago
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
arXiv·9d ago
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
arXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
arXiv·9d ago
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
arXiv·9d ago
ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
arXiv·9d ago
Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training
arXiv·9d ago
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
arXiv·9d ago
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
arXiv·9d ago
Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
arXiv·9d ago
Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification
arXiv·9d ago
Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
arXiv·9d ago
Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking
arXiv·9d ago
LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization
arXiv·9d ago
SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling
arXiv·9d ago
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
arXiv·9d ago
StreamScout: Learning When to Look Deeper for Streaming Video Understanding
arXiv·9d ago
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
arXiv·9d ago
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
arXiv·9d ago
RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
Hugging Face·9d ago
BenchMIRT: What are LLM benchmarks actually measuring?
Apple Machine Learning Research·14d ago
Agent Seer: Synthesizing Scenarios from Specification Understanding