A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
Suryaansh Jain · Rahasya Barkur · Vishal G · Ryan Rossi · Franck Dernoncourt · Jack Wang · Koustava Goswami · Nedim Lipka · +3 more
WHAT THIS RESEARCH IS ABOUT
Modern vision-language models easily generate fluent high-level captions but frequently overlook fine-grained details such as object counts, textures, materials, and spatial arrangements. While recent multi-stage pipelines can recover these missed details through iterative verification and rewriting, they suffer from substantially higher inference latency.
To solve this, the researchers introduce SimLoss, a reference-free embedding-space objective for single-pass detailed image captioning. SimLoss uses an InfoNCE contrastive loss to align the model projected hidden representations with a frozen image embedding prior to text decoding, eliminating the need for human-annotated detailed captions or synthetic multi-stage pipeline data. The framework is tested both as a fully differentiable fine-tuning model called SimLoss FFT and as a black-box reward approach called SimLoss GRPO.
The results show that embedding-space supervision can deliver the quality of multi-stage verification with the speed of single-pass generation. SimLoss FFT achieves the highest precision, nearly matches the F1 score of multi-stage methods, and runs roughly 20 times faster, while SimLoss GRPO delivers the best recall across evaluated baselines.
AI-assisted summary of the paper abstract.
KEY POINTS
- Missing visual details
Standard vision-language models often omit specific attributes, counts, materials, and spatial relations in their captions, while multi-stage fix-up pipelines are slow.
- Reference-free objective
SimLoss aligns a model hidden representation with a frozen image embedding using an InfoNCE contrastive loss before text decoding, avoiding the need for fine-grained reference captions.
- Two training variants
The method is implemented as SimLoss FFT for differentiable fine-tuning and SimLoss GRPO for black-box reward-based optimization.
- High speed and accuracy
SimLoss FFT achieves the highest precision and runs about 20 times faster than multi-stage systems, while SimLoss GRPO secures the strongest recall.