Principal investigator (ÚFAL): 
Grant id: 
101039303
Duration: 
2022-2027

NG-NLG

Next-Generation Natural Language Generation

This project aims to overcome the major hurdles that prevent current state-of-the-art models for natural language generation (NLG) from real-world deployment. While deep learning and neural networks brought considerable progress in many areas of natural language processing, neural approaches to NLG remain confined to experimental use and production NLG systems are handcrafted. The reason for this is that despite the very natural and fluent outputs of recent neural systems, neural NLG still has major drawbacks:

  1. the behavior of the systems is not transparent and hard to control (the internal representation is implicit), which leads to incorrect or even harmful outputs,
  2. the models require a lot of training data and processing power do not generalize well, and are mostly English-only.

On the other hand, handcrafted models are safe, transparent and fast, but produce less fluent outputs and are expensive to adapt to new languages and domains (topics). As a result, usefulness of NLG models in general is limited. In addition, current methods for automatic evaluation of NLG outputs are unreliable, hampering system development.

The main aims of this project, directly addressing the above drawbacks, are:

  1. Develop new approaches for NLG that combine neural approaches with explicit symbolic semantic representations, thus allowing greater control over the outputs and explicit logical inferences over the data.
  2. Introduce approaches to model compression and adaptation to make models easily portable across domains and languages.
  3. Develop reliable neural-symbolic approaches for evaluation of NLG systems.

We will test our approaches on multiple NLG applications – data-to-text generation (e.g., weather or sports reports), summarization, and dialogue response generation. For example, our approach will make it possible to deploy a new data reporting system for a given domain based on a few dozen example input-output pairs, compared to thousands needed by current methods.

Publications

2026

  • Ivan Kartáč, Mateusz Lango, Ondřej Dušek. Reasoning Gets Harder for LLMs Inside A Dialogue, in: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). [Anthology]
  • Ivan Kartáč, Kristýna Onderková, Jan Bronec, Zdeněk Kasner, Mateusz Lango, Ondřej Dušek. UFAL-CUNI at SemEval-2026 Task 11: An Efficient Modular Neuro-symbolic Method for Syllogistic Reasoning, in: Proceedings of the 20th International Workshop on Semantic Evaluation. [Anthology]
  • Zdeněk Kasner and Ondřej Dušek. AnimatedLLM: Explaining LLMs with Interactive Visualizations, in: Proceedings of the Seventh Workshop on Teaching Natural Language Processing (TeachNLP 2026). [Anthology]
  • Zdeněk Kasner, Vilém Zouhar, Patrícia Schmidtová, Ivan Kartáč, Kristýna Onderková, Ondrej Platek, Dimitra Gkatzia, Saad Mahamood, Ondřej Dušek, and Simone Balloccu. LLMs as Span Annotators: A Comparative Study of LLMs and Humans, in: Proceedings of the First Workshop on Multilingual Multicultural Evaluation. [Anthology]
  • Nalin Kumar and Ondrej Dusek. Modular Monolingual Adaptation using Pretrained Language Models, in: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track). [Anthology]
  • Kristýna Onderková. Thesis Proposal: Intentional Inference for Insight Generation, in: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). [Anthology]
  • Patrícia Schmidtová, Niyati Bafna, Seth Aycock, Gianluca Vico, Wiktor Kamzela, Kathy Hämmerl, and Vilém Zouhar. 2026. How Important is ‘Perfect’ English for Machine Translation Prompts?, in: Findings of the Association for Computational Linguistics: EACL 2026. [Anthology]
  • Patrícia Schmidtová, Saad Mahamood, and Ondřej Dušek. Never Truly Out of Fashion: A Retrospective Look at Evaluation in NLG, in: Proceedings of the 1st Symposium on Natural Language Generation Evaluations. [Anthology]
  • Patricia Schmidtova, Ondrej Dusek, Saad Mahamood. HotelCheckSpan: A Benchmark Dataset for LLM Faithfulness, in: Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). [ELRA]
  • Jędrzej Warczyński, Ondrej Dusek, and Mateusz Lango. One-step Nonautoregressive Natural Language Generation with Shortcut Flow Matching Models, in: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). [Anthology]

2025

  • Pier Felice Balestruci, Ondřej Dušek, Luca Anselma, Alessandro Mazzei. Can Large Language Models Personalize Dialogues to Generational Styles?, in: Findings of the Association for Computational Linguistics: EMNLP 2025. [Anthology]
  • Michelle Elizabeth, Morgan Veyret, Miguel Couceiro, Ondřej Dušek, Lina Maria Rojas Barahona. Exploring ReAct Prompting for Task-Oriented Dialogue: Insights and Shortcomings, in: Proceedings of the 15th International Workshop on Spoken Dialogue Systems Technology. [Anthology]
  • Wiktor Kamzela, Mateusz Lango, Ondřej Dušek. SRS-Stories: Vocabulary-constrained multilingual story generation for language learning, in: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. [Anthology]
  • Wiktor Kamzela, Mateusz Lango, Ondřej Dušek. You are an LLM teaching a smaller model everything you know: Multi-task pretraining of language models with LLM-designed study plans, in: Proceedings of the First BabyLM Workshop. [Anthology]
  • Ivan Kartáč, Mateusz Lango, Ondřej Dušek. OpeNLGauge: An Explainable Metric for NLG Evaluation with Open-Weights LLMs, in: Proceedings of the 18th International Natural Language Generation Conference. [PDF]
  • Tom Kocmi, Sweta Agrawal, Ekaterina Artemova, Eleftherios Avramidis, Eleftheria Briakou, Pinzhen Chen, Marzieh Fadaee, Markus Freitag, Roman Grundkiewicz, Yupeng Hou, Philipp Koehn, Julia Kreutzer, Saab Mansour, Stefano Perrella, Lorenzo Proietti, Parker Riley, Eduardo Sánchez, Patricia Schmidtova, Mariya Shmatova, Vilém Zouhar. Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation, in: Proceedings of the Tenth Conference on Machine Translation. [Anthology
  • Nalin Kumar, Mateusz Lango, and Ondrej Dusek. Pretraining Language Models with LoRA and Artificial Languages, in: Proceedings of the First BabyLM Workshop. [Anthology
  • Nalin Kumar. Thesis Proposal: Efficient Methods for Natural Language Generation/Understanding Systems, in: The 14th International Joint Conference on Natural Language Processing and The 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. [Anthology
  • Mateusz Lango, Ondřej Dušek. LLM Agents Implement an NLG System from Scratch: Building Interpretable Rule-Based RDF-to-Text Generators, in: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. [Anthology]
  • Karen Jia-Hui Li, Simone Balloccu, Ondřej Dušek, Ehud Reiter. When LLMs Can’t Help: Real-World Evaluation of LLMs in Nutrition, in: Proceedings of the 18th International Natural Language Generation Conference. [PDF]
  • Sourabrata Mukherjee, Atul Kr. Ojha, John McCrae, Ondřej Dušek. Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?, in: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop). [Anthology]
  • Kristýna Onderková, Mateusz Lango, Patrícia Schmidtová, Ondřej Dušek. ReproHum #0669-08: Reproducing Sentiment Transfer Evaluation, in: Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²). [Anthology]
  • Kristýna Onderková, Ondřej Plátek, Zdeněk Kasner, Ondřej Dušek. FreshTab: Sourcing Fresh Data for Table-to-Text Generation Evaluation, in: Proceedings of the 18th International Natural Language Generation Conference. [PDF]
  • Patrícia Schmidtová, Eduardo Calò, Simone Balloccu, Dimitra Gkatzia, Rudali Huidrom, Mateusz Lango, Fahime Same, Vilém Zouhar, Saad Mahamood, Ondřej Dušek. Do My Eyes Deceive Me? A Survey of Human Evaluations of Hallucinations in NLG, in: Proceedings of the 18th International Natural Language Generation Conference. [PDF]
  • Patrícia Schmidtová, Ondřej Dušek, Saad Mahamood. Real-World Summarization: When Evaluation Reaches Its Limits, in: Findings of the Association for Computational Linguistics: EMNLP 2025. [Anthology]
  • Alex Terentowicz, Mateusz Lango, Ondřej Dušek. How (un)faithful are explainable LLM-based NLG metrics?, in: Proceedings of the 18th International Natural Language Generation Conference. [PDF]

2024

  • Adam Wojciechowski, Mateusz Lango, Ondrej Dusek. Faithful and Plausible Natural Language Explanations for Image Classification: A Pipeline Approach, in: Findings of the Association for Computational Linguistics: EMNLP 2024. [Anthology][Github][Poster]
  • Simone Balloccu, Ehud Reiter, Karen Jia-Hui Li, Rafael Sargsyan, Vivek Kumar, Diego Reforgiato, Daniele Riboni, Ondrej Dusek. Ask the experts: sourcing a high-quality nutrition counseling dataset through Human-AI collaboration, in: Findings of the Association for Computational Linguistics: EMNLP 2024. [Anthology][Github][Poster]
  • Sourabrata Mukherjee, Atul Kr. Ojha, Akanksha Bansal, Deepak Alok, John P. McCrae, Ondrej Dusek. Multilingual Text Style Transfer: Datasets & Models for Indian Languages, in: Proceedings of the 17th International Natural Language Generation Conference. [Anthology][Github code][Github data][Poster]
  • Sourabrata Mukherjee, Atul Kr. Ojha, Ondrej Dusek. Are Large Language Models Actually Good at Text Style Transfer?, in: Proceedings of the 17th International Natural Language Generation Conference. [Anthology][Slides][Github]
  • Patricia Schmidtova, Saad Mahamood, Simone Balloccu, Ondrej Dusek, Albert Gatt, Dimitra Gkatzia, David M. Howcroft, Ondrej Platek, Adarsa Sivaprasad. Automatic Metrics in Natural Language Generation: A survey of Current Evaluation Practices, in: Proceedings of the 17th International Natural Language Generation Conference. [Anthology][Github][Slides]
  • Jędrzej Warczyński, Mateusz Lango, Ondrej Dusek. Leveraging Large Language Models for Building Interpretable Rule-Based Data-to-Text Systems, in: Proceedings of the 17th International Natural Language Generation Conference. [Anthology][Github][Slides]
  • Zdeněk Kasner, Ondrej Platek, Patricia Schmidtova, Simone Balloccu, Ondrej Dusek. factgenie: A Framework for Span-based Evaluation of Generated Texts, in: Proceedings of the 17th International Natural Language Generation Conference: System Demonstrations. [Anthology][Github][PyPI][Poster]
  • Zdeněk Kasner, Ondrej Dusek. Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). [Anthology][Website][Github][Poster][Slides]
  • Xing Han Lu, Zdeněk Kasner, Siva Reddy. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue, in: Proceedings of the 41st International Conference on Machine Learning. [Paper text][Website]
  • Nalin Kumar, Ondrej Dusek. LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems, in: Findings of the Association for Computational Linguistics: NAACL 2024. [Anthology][Github][Poster]
  • Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, Ondrej Dusek. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs, in: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). [Anthology][Website][Poster][Slides]
  • Mateusz Lango, Patrícia Schmidtová, Simone Balloccu, Ondřej Dušek. ReproHum #0043-4: Evaluating Summarization Models: Investigating the Impact of Education and Language Proficiency on Reproducibility, in: The 4th Workshop on Human Evaluation of NLP Systems (HumEval’24). [Anthology][Slides][Poster]

2023

  • Zdeněk Kasner, Ioannis Konstas, Ondrej Dusek. Mind the Labels: Describing Relations in Knowledge Graphs With Pretrained Models, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. [Anthology][Github] [Poster]
  • Zixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, Daniele Riboni. Are Experts Needed? On Human Evaluation of Counselling Reflection Generation, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). [Anthology]
  • Mateusz Lango, Ondrej Dusek. Critic-Driven Decoding for Mitigating Hallucinations in Data-to-text Generation, in: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. [Anthology][Github][Slides]
  • Zdeněk Kasner, Ekaterina Garanina, Ondrej Platek, Ondrej Dusek. TabGenie: A Toolkit for Table-to-Text Generation, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). [Anthology][Demo] [PyPi] [Poster]
  • Vojtěch Hudeček, Ondrej Dusek. Are Large Language Models All You Need for Task-Oriented Dialogue?, in: Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. [Anthology][Github][Slides]
  • Sourabrata Mukherjee, Ondrej Dusek. Leveraging Low-resource Parallel Data for Text Style Transfer, in: Proceedings of the 16th International Natural Language Generation Conference. [Anthology][Github][Slides]
  • Saad Obaid ul Islam, Iza Škrjanec, Ondrej Dusek, Vera Demberg. Tackling Hallucinations in Neural Chart Summarization, in: Proceedings of the 16th International Natural Language Generation Conference. [Anthology][Github][Poster]
  • Mateusz Woźny, Mateusz Lango. Generating clickbait spoilers with an ensemble of large language models, in: Proceedings of the 16th International Natural Language Generation Conference. [Anthology]
  • Jakub Raczyński, Mateusz Lango, Jerzy Stefanowski. The Problem of Coherence in Natural Language Explanations of Recommendations, in: Proceedings of the 26th European Conference on Artificial Intelligence (ECAI 2023). [Paper text]
  • Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Stephanie Schoch, Craig Thomson, Luou Wen. Barriers and enabling factors for error analysis in NLG research, in: Northern European Journal of Language Technology. [Paper text]
  • Ondřej Plátek, Ondrej Dusek. MooseNet: A Trainable Metric for Synthesized Speech with a PLDA Module, in: 12th ISCA Speech Synthesis Workshop (SSW2023). [Paper text][Slides][Github]
  • Ondrej Platek, Mateusz Lango, Ondrej Dusek. With a Little Help from the Authors: Reproducing Human Evaluation of an MT Error Detector, in: Proceedings of the 3rd Workshop on Human Evaluation of NLP Systems. [Anthology][Github][Slides]
  • Sourabrata Mukherjee, Akanksha Bansal, Pritha Majumdar, Atul Kr. Ojha, Ondřej Dušek. Low-Resource Text Style Transfer for Bangla: Data & Models, in: Proceedings of the First Workshop on Bangla Language Processing (BLP-2023). [Anthology][Slides][Github code][Github data]
  • Sourabrata Mukherjee, Vojtěch Hudeček, Ondřej Dušek. Polite Chatbot: A Text Style Transfer Application, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop. [Anthology][Github][Poster][Slides]
  • Nalin Kumar, Saad Obaid Ul Islam, Ondrej Dusek. Better Translation + Split and Generate for Multilingual RDF-to-Text (WebNLG 2023), in: Proceedings of the Workshop on Multimodal, Multilingual Natural Language Generation and Multilingual WebNLG Challenge (MM-NLG 2023). [Anthology][Github][Slides]
  • Ondřej Plátek, Vojtech Hudecek, Patricia Schmidtova, Mateusz Lango, Ondrej Dusek. Three Ways of Using Large Language Models to Evaluate Chat, in: Proceedings of The Eleventh Dialog System Technology Challenge. [Anthology][Poster][Github]
  • František Trebuňa, Ondrej Dusek. VisuaLLM: Easy Web-based Visualization for Neural Language Generation, in: Proceedings of the 16th International Natural Language Generation Conference: System Demonstrations. [Anthology][Poster][Github]
  • Patricia Schmidtova. Semantic Accuracy in Natural Language Generation: A Thesis Proposal, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop). [Anthology]

2022

  • Zdeněk Kasner, Ondřej Dušek. Neural Pipeline for Zero-Shot Data-to-Text Generation, in: ACL [Anthology] [Github] [Poster]
  • Tomáš Nekvinda, Ondřej Dušek. AARGH! End-to-end Retrieval-Generation for Task-Oriented Dialog, in: SIGdial. [arXiv] [video] [Github]
  • Sourabrata Mukherjee, Zdeněk Kasner, Ondřej Dušek. Balancing the Style-Content Trade-Off in Sentiment Transfer Using Polarity-Aware Denoising, in: Text, Speech and Dialogue. [SpringerLink]
  • Rudali Huidrom, Ondřej Dušek, Zdeněk Kasner, Thiago Castro Ferreira, Anya Belz. Two Reproductions of a Human-Assessed Comparative Evaluation of a Semantic Error Detection System, in: INLG GenChal [Anthology]
  • Vojtěch Hudeček, Ondřej Dušek. Learning Interpretable Latent Dialogue Actions With Less Supervision, in: AACL-IJCNLP [arXiv] [Github]

Media

  • Ondřej Dušek: Problémy dnešních generátorů jazyka. Vesmír 101, Sep 2022 [URL]
  • Natural language generation research wins ERC grant. Forum Charles University Magazine. Jan 2022 [URL]