GRIDINDEX TOPIC

Language Models

40 tracked items

RESEARCHEPFL News·9d ago
An AI capable of doubt can optimize scientific discovery
RESEARCHarXiv·9d ago
The Emergent Symbolic Structure of Artificial Neural Networks
RESEARCHarXiv·9d ago
Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
RESEARCHarXiv·9d ago
Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding
RESEARCHarXiv·9d ago
A Cognitive Architecture for Shared Autonomy in AUV Operations
RESEARCHarXiv·9d ago
LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark
RESEARCHarXiv·9d ago
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
RESEARCHarXiv·9d ago
Incremental Risk Assessment of Progressive Elder Financial Scams via Instruction-Tuned Small Language Models
RESEARCHarXiv·9d ago
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation
RESEARCHarXiv·9d ago
ViTAL-X: Video-Text Alignment with Cross-Modal Temporal Edits
RESEARCHarXiv·9d ago
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
RESEARCHarXiv·9d ago
Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models
RESEARCHarXiv·9d ago
Beyond Blind Compliance: Benchmarking Task Verification in OCR Reasoning
RESEARCHarXiv·9d ago
Less Is More: Balancing Positive and Negative Space in Visual Concept Blending
RESEARCHarXiv·9d ago
Authority Bias in Conversational Search Engines for Academic Paper Recommendation
RESEARCHarXiv·9d ago
Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning
RESEARCHarXiv·9d ago
The Potential of Haptic Foundation Models
RESEARCHarXiv·9d ago
Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
RESEARCHarXiv·9d ago
Good Memory Has ECC: Evaluating the Memory of Vision-Language Models Beyond Accuracy
RESEARCHarXiv·9d ago
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
RESEARCHarXiv·9d ago
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling
RESEARCHarXiv·9d ago
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
RESEARCHarXiv·9d ago
CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language
RESEARCHarXiv·9d ago
Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy
RESEARCHarXiv·9d ago
Medical Causal Hypothesis Verification with Large Language Models
RESEARCHarXiv·9d ago
Generative artificial intelligence for reliable mechanistic reasoning for corrosion
RESEARCHarXiv·9d ago
QTEA: Ternary LLMs with Sparse Residual Salient Weight and By-Column Optimization
RESEARCHarXiv·9d ago
Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment
RESEARCHarXiv·9d ago
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
RESEARCHarXiv·9d ago
Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs
RESEARCHarXiv·9d ago
LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts
RESEARCHarXiv·9d ago
Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs
RESEARCHarXiv·9d ago
AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
RESEARCHarXiv·9d ago
Asymmetries in Spontaneous and Instructed Deception
RESEARCHarXiv·9d ago
OpenAgentFlow: Enabling System-Wide Safety Boundaries for Heterogeneous AI Agent Fleets
RESEARCHarXiv·9d ago
MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts
RESEARCHmedRxiv·9d ago
Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology
RESEARCHmedRxiv·9d ago
Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation
VentureBeat·9d ago
Frontier models can recover up to 65% of facts they can't directly recall — just by thinking longer
VentureBeat·9d ago
Anthropic's Claude Fable 5.1 and Mythos 5.1 arrive with a 75% cost reduction for Fable cache reads