Accuracy
38 tracked items
The Guardian·8d ago
Australian science prize awarded to researcher studying memory accuracy of domestic violence survivors
bioRxiv·8d ago
Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation
NVIDIA Developer·8d ago
Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
STAT·9d ago
STAT+: How a former ARPA-H director’s startup is tackling AI’s ‘dumb problems’
Tom's Hardware·9d ago
Researchers build a $7 smartphone clip-on that spots hidden cameras — AI and dynamic LED grid deliver 94% accuracy
The Decoder·9d ago
Google Gemini's new agent-based video analysis cuts token usage by up to 88 percent
arXiv·9d ago
Medical Causal Hypothesis Verification with Large Language Models
arXiv·9d ago
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
arXiv·9d ago
From Multi-Modal Paths to Executable Trajectories: A Trajectory Planning Framework for 4WIS Robots
arXiv·9d ago
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
arXiv·9d ago
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
arXiv·9d ago
Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
arXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
arXiv·9d ago
Do General NLP Embeddings Capture Ontological Reasoning?
arXiv·9d ago
A Degradation-Tolerance Benchmark for Camera-Only End-to-End Driving
arXiv·9d ago
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
arXiv·9d ago
DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
arXiv·9d ago
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
arXiv·9d ago
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
arXiv·9d ago
Foundation models for electricity price forecasting and battery arbitrage: Can they replace market-specific forecasting models?
arXiv·9d ago
Coding What Matters: A Semantic-Aware Memory Interface for Energy-Efficient Perception in Autonomous Vehicles
arXiv·9d ago
Different representation learning objectives recover distinct latent structures from the same psychometric data
arXiv·9d ago
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
arXiv·9d ago
Brain-Language-Action (BLA) Models: Language-Conditioned EEG for Robotics Control
arXiv·9d ago
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
arXiv·9d ago
Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment
arXiv·9d ago
Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure
arXiv·9d ago
Long-Horizon State Tracking in LLMs: Executing MD5 through a Deep Sequence of Dependent Tool Calls
arXiv·9d ago
Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers
arXiv·9d ago
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
arXiv·9d ago
MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts
arXiv·9d ago
SlideMix: Enhancing Whole Slide Image Analysis via Multimodal Shuffling
medRxiv·9d ago
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference
medRxiv·9d ago
The accuracy of urine-based mycobacterial antigens to detect childhood tuberculosis using an ultrasensitive immunoassay
medRxiv·9d ago
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation
medRxiv·9d ago
Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology
Towards AI·10d ago
Beyond a Single Model: Mastering Ensemble Learning in ML
AWS Machine Learning·13d ago
How Decathlon runs demand forecasting at scale with Chronos-2