Evaluating
11 tracked items
arXiv·9d ago
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
arXiv·9d ago
Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations
arXiv·9d ago
AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning
arXiv·9d ago
When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
arXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
arXiv·9d ago
Do General NLP Embeddings Capture Ontological Reasoning?
arXiv·9d ago
Medical Causal Hypothesis Verification with Large Language Models
arXiv·9d ago
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
medRxiv·9d ago
Evaluating Mean Platelet Volume in relation to Disease Severity in Paediatric Sickle Cell Anaemia: A Cross-Sectional Study in Kwara State, North-Central Nigeria
medRxiv·9d ago
A Randomized Controlled Trial Evaluating a Community-Based, Family Network Heart Health Intervention - the SERVE OC Trial: Design, Rationale and Baseline Findings
Apple Machine Learning Research·14d ago
Agent Seer: Synthesizing Scenarios from Specification Understanding