News

Successful submissions to ACL 2026

09/07/2026

At the annual meeting of the Association for Computational Linguistics (ACL), ATHENE researcher Prof. Dr. Iryna Gurevych and her colleagues presented numerous research papers. A paper co-authored by ATHENE researchers Prof. Dr. Anna Rohrbach and Prof. Dr. Marcus Rohrbach was also honored with the Best Resource Paper Award. This highly regarded conference brings together researchers and experts in the field of computational linguistics and language technology (natural language processing) to discuss the latest developments in this area.

Accepted papers are: 

Protecting multimodal large language models against misleading visualizations
Authors: Jonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna Gurevych
Charts help make data easier to understand, for example in news reporting or science communication. But some charts are deliberately designed to mislead, for instance through truncated or inverted axes. Such tricks lead people to draw incorrect conclusions. The researchers examined whether multimodal AI language models — AI systems that understand text and images together — can be fooled the same way. They tested 19 such models, including well-known commercial systems. The result: on misleading charts, model accuracy drops to nearly the level of random guessing. The researchers then tested six methods to fix this weakness. Two methods prove effective: first, having the model extract a data table from the chart and answer based on that table; second, having the model redraw the chart from the extracted data in an undistorted way. Both methods improve accuracy by up to 19.6 percentage points. The findings show that not only humans, but also AI systems, need to be specifically protected against manipulated charts, for example to help curb the spread of disinformation. This work was conducted as part of the ATHENE project Safeguarding LLMs against Misleading Evidence Attack (SafeLLMs), which is part of the ATHENE research area Reliable and Verifiable Information through Secure Media (REVISE).
PDF:https://arxiv.org/pdf/2502.20503

Is this chart lying to me? Automating the detection of misleading visualizations
Authors: Jonathan Tonglet, Jan Zimny, Tinne Tuytelaars, Iryna Gurevych

Manipulated charts are a common way to spread false information online and on social media. Until now, large, freely available datasets to train and test AI systems for detecting such charts have been missing. The researchers therefore introduce two new datasets. The first, Misviz, contains 2,604 real-world charts, each labeled with one of twelve known types of deceptive design. The second, Misviz-synth, includes 57,665 artificially generated charts based on real-world data tables. Using these datasets, the researchers compared three detection approaches: multimodal AI language models, rule-based checking programs, and specially trained image classifiers. The result: the task remains difficult for all approaches. AI language models perform best on real-world charts, while the other methods have an advantage on artificially generated ones. The researchers make both datasets and the accompanying code is freely available. This supports future research and can help protect readers from manipulated charts, as well as alert chart creators to unintentional design errors. 
This work was also conducted as part of the ATHENE project Safeguarding LLMs against Misleading Evidence Attack (SafeLLMs), which is part of the ATHENE research area Reliable and Verifiable Information through Secure Media (REVISE).
PDF: https://arxiv.org/pdf/2508.21675

Patches of Nonlinearity: Instruction Vectors in Large Language Models
Authors: Irina Bigoulaeva, Jonas Rohweder, Subhabrata Dutta, Iryna Gurevych
Large AI language models are usually trained through additional fine-tuning to follow instructions. Until now, however, little was known about how these models process such instructions internally. The researchers therefore investigated where in the model an instruction is stored and how it is further processed. They found that the model forms a compact internal representation right after reading the instruction, similar to a kind of short summary. These representations for different tasks can be clearly distinguished from one another. However, the way they work together follows surprisingly complex, non-linear patterns: several such representations act more strongly together than their individual contributions would suggest. This challenges a common assumption in AI research, namely that individual internal building blocks of AI models can be understood independently of one another. The researchers therefore developed a new analysis method that accounts for these complex interactions. Using it, they show that instruction representations act like a switch that determines which internal processing pathways the model uses for a given task. These findings help make AI language models better understood, more reliably verifiable, and safer for practical use.
The work was conducted as part of the ATHENE project SecLLM: Security in Large Language Models, a project of the ATHENE research area Security and Privacy in Artificial Intelligence (SenPAI).
PDF: https://arxiv.org/pdf/2602.07930.

SCICOQA: Quality Assurance for Scientific Paper–Code Alignment
Authors: Tim Baumgärtner and Iryna Gurevych
Scientific studies often publish not only a text but also the accompanying program code. However, the text and code do not always match: the code sometimes implements something different from what the paper describes. Such mismatches threaten the reproducibility of research findings. The problem is growing because "AI scientists" increasingly generate studies and code on their own, making it impossible for humans to check everything. The researchers therefore introduce SCICOQA, the first benchmark dataset for testing how well AI language models detect such mismatches. It contains 635 examples: 92 real-world cases and 543 artificially generated cases from various scientific fields. The result is sobering: even the best AI models tested detect only 46.7 percent of real mismatches between paper and code. Models find it especially hard to spot things that appear in the code but are missing from the paper. To support the development of better, automated quality-control tools for science, the researchers publish the dataset openly. It is publicly availabale via https://ukplab.github.io/scicoqa/. The work was conducted as part of the ATHENE mission project SecureCoder .
PDF: https://arxiv.org/pdf/2601.12910.

Other papers co-authored by ATHENE researchers that were accepted for the conference include:

In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
Authors: Hiba Arnaout, Noy Sternlicht, Tom Hope, Iryna Gurevych
PDF: https://arxiv.org/pdf/2505.14838

Author-in-the-Loop Response Generation and Evaluation: Integrating Author Expertise and Intent in Responses to Peer Review
Authors: Qian Ruan, Iryna Gurevych
PDF: https://arxiv.org/pdf/2602.11173

Reward Modeling for Scientific Writing Evaluation
Authors: Furkan Sahinuç  Subhabrata Dutta, Iryna Gurevych
PDF:  https://arxiv.org/pdf/2601.11374

Responsible Evaluation of AI for Mental Health
Authors: Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych
PDF: https://arxiv.org/pdf/2602.00065

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
Authors: Mark Rothermel, Marcus Kornmann, Marcus Rohrbach, Anna Rohrbach
PDF: https://arxiv.org/pdf/2601.08611
The paper was awarded the Best Resource Paper Award.

TACL Paper:
Can LLMs Automate Fact-Checking Article Writing?
Authors: Dhruv Sahnan, David Corney, Irene Larraz, Giovanni Zagni, Ruben Miguez, Zhuohan Xie, Iryna Gurevych, Elizabeth Churchill, Tanmoy Chakrabortyl, Preslav Nakov 
PDF: https://arxiv.org/pdf/2503.17684

ACL 2026 took place July 2–7 in San Diego, USA.

show all news