GRIDINDEX TOPIC

Benchmark

40 tracked items

Tom's Hardware·8d ago
Benchmarking 31 different CPUs in Onimusha: Way of the Sword — X3D beats flagships by 10%, 270K Plus falls behind Raptor Lake Refresh
RESEARCHbioRxiv·8d ago
Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation
KDnuggets·9d ago
This Python Library Can Run Pandas Workloads Up to 20x Faster
RESEARCHarXiv·9d ago
LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
RESEARCHarXiv·9d ago
I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
RESEARCHarXiv·9d ago
AnyWorld: Factorized Egocentric World Models for Cross-Embodiment Generalization
RESEARCHarXiv·9d ago
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
RESEARCHarXiv·9d ago
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
RESEARCHarXiv·9d ago
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
RESEARCHarXiv·9d ago
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
RESEARCHarXiv·9d ago
UI-Venus-2 Technical Report
RESEARCHarXiv·9d ago
ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
RESEARCHarXiv·9d ago
ReDeck: Step-Level Render-Grounded Refinement for Document-to-Slide Generation
RESEARCHarXiv·9d ago
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
RESEARCHarXiv·9d ago
Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
RESEARCHarXiv·9d ago
The Potential of Haptic Foundation Models
RESEARCHarXiv·9d ago
Multi-Group Pipe Routing under Permanent Geometric Occupancy: Problem, Benchmark, and Classical Baselines
RESEARCHarXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
RESEARCHarXiv·9d ago
Do General NLP Embeddings Capture Ontological Reasoning?
RESEARCHarXiv·9d ago
A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving
RESEARCHarXiv·9d ago
Auditing Harness Tampering in Self-Improving Agents
RESEARCHarXiv·9d ago
Task-Specific Prompt with Global Context for Multi-Task Graph Pre-Training
RESEARCHarXiv·9d ago
DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
RESEARCHarXiv·9d ago
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
RESEARCHarXiv·9d ago
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
RESEARCHarXiv·9d ago
Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
RESEARCHarXiv·9d ago
Local Reference Geometry Residual Augmentation for Imbalanced Time Series Classification
RESEARCHarXiv·9d ago
Elite-Weighted Supervised Fine-tuning for Goal-Directed Molecular Optimization
RESEARCHarXiv·9d ago
Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
RESEARCHarXiv·9d ago
Geometry-aware Latent Autoregressive Generative Model for PDEs in Complex Domains
RESEARCHarXiv·9d ago
Neural means and kernel corrections for operator learning
RESEARCHarXiv·9d ago
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
RESEARCHarXiv·9d ago
AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
RESEARCHarXiv·9d ago
Bridging Lexical Divergence: LLM-Assisted, Cost-Efficient, Zero-shot Scientific Entity Linking
RESEARCHarXiv·9d ago
LOOMSUM:Weaving Quantitative and Narrative Evidence for Faithful Long Text-Table Summarization
RESEARCHarXiv·9d ago
SGE: Semantically-Guided Exploration for Unstructured Environments via Image-Space Waypoint Sampling
RESEARCHarXiv·9d ago
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
RESEARCHarXiv·9d ago
Discrete-Time MDP Modeling for Multi-Item Capacitated Lot Sizing with Stochastic Demand Timing
RESEARCHarXiv·9d ago
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
RESEARCHarXiv·9d ago
StreamScout: Learning When to Look Deeper for Streaming Video Understanding