Skip to main navigation Skip to search Skip to main content

Measure only what is measurable: Towards conversation requirements for evaluating task-oriented dialogue systems

Research output: Chapter in Book/Report/Conference proceedingConference contributionScientificpeer-review

5 Downloads (Pure)

Abstract

Chatbots for customer service have been widely studied in many different fields, ranging from Natural Language Processing (NLP) to Communication Science. These fields have developed different evaluation practices to assess chatbot performance (e.g., fluency, task success) and to measure the impact of chatbot usage on the user's perception of the organisation controlling the chatbot (e.g., brand attitude) as well as their willingness to enter a business transaction or to continue to use the chatbot in the future (i.e., purchase intention, reuse intention). While NLP researchers have developed many automatic measures of success, other fields mainly use questionnaires to compare different chatbots. This paper explores the extent to which we can bridge the gap between the two, and proposes a research agenda to further explore this question.
Original languageEnglish
Title of host publicationProceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
EditorsKaustubh Dhole, Miruna Clinciu
Place of PublicationVienna, Austria and virtual meeting
PublisherAssociation for Computational Linguistics
Pages231-238
Number of pages8
ISBN (Print)979-8-89176-261-9
Publication statusPublished - 1 Jul 2025
Event4th Workshop on Generation, Evaluation, and Metrics - The Austria Center Vienna (hybrid), Vienna, Austria
Duration: 31 Jul 20251 Aug 2025
https://gem-benchmark.com/workshop

Conference

Conference4th Workshop on Generation, Evaluation, and Metrics
Abbreviated titleGEM 2025
Country/TerritoryAustria
CityVienna
Period31/07/251/08/25
Internet address

Keywords

  • natural language generation
  • evaluation metrics
  • prompt engineering
  • benchmarking
  • model assessment

Fingerprint

Dive into the research topics of 'Measure only what is measurable: Towards conversation requirements for evaluating task-oriented dialogue systems'. Together they form a unique fingerprint.

Cite this