Inference
32 tracked items
CNBC·8d ago
How Equinix has found a niche in the multitrillion-dollar AI data center boom
AWS Machine Learning·8d ago
Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference
NVIDIA Developer·8d ago
Co-Designing AI Models Using Speculative Decoding for Faster LLM Inference
Tom's Hardware·9d ago
Startups want to rent your idle gaming PC for AI tasks — Startups pitch an 'Airbnb for AI inference,' but profitability remains unproven
arXiv·9d ago
Beyond Language Priors: Diagnosing and Fixing Visual-Origin Hallucinations in Multimodal LLM
arXiv·9d ago
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
arXiv·9d ago
Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
arXiv·9d ago
Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation
arXiv·9d ago
Instance-Guided Report Anchoring for Text-Free 3D Abnormality Segmentation in Chest CT
arXiv·9d ago
Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance
arXiv·9d ago
DISTAL: Distillation and Self-Supervised Pretraining for Structure-Agnostic Materials Property Prediction
arXiv·9d ago
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
arXiv·9d ago
OCGQuant: Outlier-Companion Grouping for NVFP4 Quantization
arXiv·9d ago
Flawed in Nature, Perfect through Evolution
arXiv·9d ago
AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
arXiv·9d ago
When Prediction Error Is Not Enough: Evaluating Nuisance-Function Prediction for Causal Estimation
arXiv·9d ago
StreamScout: Learning When to Look Deeper for Streaming Video Understanding
arXiv·9d ago
IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
arXiv·9d ago
Learning What to Retain: Gated-Memory Routing for Efficient Collaboration in Multi-Agent LLM Systems
medRxiv·9d ago
Can Dental AI Really Beat Dentists? DentalPair-Cert for Rigorous AI-Dentist Inference
Nature·9d ago
Robust inference and correlates from genetic associations with personality
NVIDIA Developer·10d ago
How to Size GPUs for AI Inference and TCO Without Overspending
The Information·10d ago
Wafer, An Inference Provider That Uses Non-Nvidia Chips, Lands Acquisition Offers and $200 Million-Plus Valuation
IEEE Spectrum AI·10d ago
Cash In on the AI Boom by Renting Out Your Spare Compute
KDnuggets·11d ago
Speed Up LLM Inference with DSpark Speculative Decoding
NVIDIA Developer·13d ago
Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect
AWS Machine Learning·13d ago
How Decathlon runs demand forecasting at scale with Chronos-2
AWS Machine Learning·13d ago
Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components
AI Business·14d ago
Z.AI's Use of Chinese Chips for New Model is About Optimization
AWS Machine Learning·14d ago
Introducing OpenAI models on Amazon Bedrock for in-country inferencing in India
AWS Machine Learning·14d ago
Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2
AI Business·15d ago
Qwen 3.8 Flash-Next is Cheap, But There Are Complicating Factors