Skip to main navigation Skip to search Skip to main content

Tailoring LLM-generated image captions to user needs

Research output: Contribution to journalArticleScientificpeer-review

5 Downloads (Pure)

Abstract

One of the original motivations for the development of image captioning systems is to make visual content accessible for people who are blind or visually impaired. What seemed like a huge challenge fifteen years ago, has now made it into consumer products: large language models such as ChatGPT are seemingly able to describe images in fluent natural language. But it is still unclear to what extent the generated descriptions actually match user needs. This study investigates the quality of LLM-generated image descriptions in the context of Dutch news articles. We operationalise output quality based on earlier user studies and existing image description guidelines, and present an extensive evaluation protocol that may be used in future research to assess the quality of automatically generated image descriptions.
Original languageEnglish
Pages (from-to)165-191
Number of pages27
JournalComputational Linguistics in the Netherlands Journal
Volume15
Publication statusPublished - 2026

Fingerprint

Dive into the research topics of 'Tailoring LLM-generated image captions to user needs'. Together they form a unique fingerprint.

Cite this