GRIDINDEX TOPIC

Benchmarks

22 tracked items

RESEARCHarXiv·9d ago
A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving
RESEARCHarXiv·9d ago
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
RESEARCHarXiv·9d ago
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
RESEARCHarXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
RESEARCHarXiv·9d ago
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
RESEARCHarXiv·9d ago
ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
RESEARCHarXiv·9d ago
Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training
RESEARCHarXiv·9d ago
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
RESEARCHarXiv·9d ago
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
RESEARCHarXiv·9d ago
Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
RESEARCHarXiv·9d ago
Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification
RESEARCHarXiv·9d ago
Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
RESEARCHarXiv·9d ago
Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking
RESEARCHarXiv·9d ago
LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization
RESEARCHarXiv·9d ago
SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling
RESEARCHarXiv·9d ago
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
RESEARCHarXiv·9d ago
StreamScout: Learning When to Look Deeper for Streaming Video Understanding
RESEARCHarXiv·9d ago
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
RESEARCHarXiv·9d ago
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
RESEARCHarXiv·9d ago
RoboPhys-3D: A Comprehensive Embodied World Model Evaluation via 3D Reconstruction
Hugging Face·9d ago
BenchMIRT: What are LLM benchmarks actually measuring?
Apple Machine Learning Research·14d ago
Agent Seer: Synthesizing Scenarios from Specification Understanding