From Noisy Text to Real-World Applications of LLMs
The WNUT workshop on language by people. We focus on language as it occurs in the real world, from noisy text to real-world applications of LLMs.
Shared Task
This year, we host MultiLexNorm2026: a shared task on multi-lingual lexical normalization with a focus on non Indo-european languages. After the success of our first MultiLexNorm shared task held in 2021, we have extended our benchmark to more varied languages. More information about MultiLexNorm2026.
Invited Speakers
Important Dates
| Date | Event |
|---|---|
| July 25th | Submission Deadline (anytime on earth; dual-submission allowed) |
| August 22nd | ARR Commitment Date |
| August 25th | Acceptance Notification |
| September 6th | Camera-Ready Deadline |
| September 11th | EMNLP 2026 Findings Deadline |
| October 28th | Workshop Day |
Workshop Schedule
| Time | Name |
|---|---|
| 09:00 - 09:30 | Welcome |
| 09:30 - 10:30 | Keynote 1 (TBA) |
| 10:30 - 11:00 | Coffee Break & Networking |
| 11:00 - 12:30 | Oral Presentation |
| 12:30 - 14:00 | Lunch |
| 14:00 - 15:00 | Keynote 2 (TBA) |
| 15:00 - 16:00 | Poster Presentation |
| 15:30 - 16:00 | Coffee Break & Networking |
| 16:00 - 17:00 | Keynote 3 (TBA) |
| 17:00 - 17:30 | Closing & Next Steps |
Accepted Papers
26 accepted papers, listed alphabetically by title.
-
A Precision-First Whitespace Restoration Pipeline for Noisy Urdu Tafseer Text
-
Are More Emotion Theories Better? Joint Prompting Improves Emotion Categorization Robustness
-
Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text
-
CrisisKD: Five-Stage Knowledge Distillation for Aspect-Level Sentiment and Emotion Analysis in Crisis Discourse
-
Discourse-Structural Noise in LLM-Generated Explanations
-
Do LLMs Give Consistent Opinions? Evaluating Response Reliability Under Varying Likert-Scale Formulations in Survey-Style MCQA
-
Evaluating Token Probabilities as Confidence Estimators in Human Behavior Simulations: A Case Study in Luxembourgish Next Like Prediction
-
From Errors to Precision: A Taxonomy-Guided Approach for LLM Flood Cause Extraction
-
GeNERscore: Semantic Alignment Metrics for Generative Named Entity Recognition
-
How Noisy is Automatic Annotation? Characterizing Label Noise from Generative Models in Multi-Dimensional Spanish Subjectivity Annotation
-
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
-
LLMs are Easily Startled: Fragility Testing with Lookalike Character Noise
-
Multi-author style detection with stylistic embeddings
-
NERQual: Evaluating the Robustness of Named Entity Recognition Models to Data Quality Issues
-
Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations
-
Personalized Emotion Analysis via Prompting Large Language Models with Writer Information
-
Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer
-
QueryDiagramEval: Evaluating and Benchmarking LLM Diagram Generation under Noisy Conditions
-
Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts
-
RSPC: A Benchmark for Modeling Stress and Psychiatric Conditions in Digitally Mediated Relationships using Psychiatrist Annotations
-
Spoiler Court: Interpretable Spoiler Detection via Adversarial Multi-Agent LLM Judgement
-
The Curse of Multilinguality in Lexical Normalization
-
Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts
-
TurEngMix: A Text Corpus and Benchmark for Turkish-English Code-Mixed Language Identification and Named Entity Recognition
-
When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text
-
Where Does Noise Robustness Live? Evidence from Ablation, Role Swaps, and Cross-Lingual Transfer in Code-Mixed Banking NLU
Call for Papers
We seek submissions of long and short papers on original and unpublished work (same page limit as the EMNLP 2026 main conference). All accepted submissions will be presented as talks and/or posters at the workshop, following the EMNLP 2026 main conference.
We welcome submissions addressing (but not limited to) the following areas:
Classical NLP Tasks on Noisy Text
- NLP of noisy text, e.g., POS and NER tagging, parsing
- Text normalization and error correction
- Paraphrase identification and semantic similarity of short text or noisy text
- Extracting user demographics, profiles, and major life events
- Machine translation and multilingual NLP over noisy text
- Information extraction from noisy text and event extraction
- Colloquial language, e.g., idiom detection
- Domain adaptation to user-generated text
- Detecting rumors, contradictory information, sarcasm, and humor on social media
- Sentiment analysis
- Temporal aspects of user-generated content (resolving time expressions, concept drift, etc.)
- Representing and mining language variation in user-generated content
LLMs and Noisy Text
- Robustness of LLMs to noisy, ungrammatical, or informal input
- Training and fine-tuning of LLMs on user-generated text
- Evaluation methodologies for LLMs on noisy text
- Domain adaptation of LLMs to specific user-generated text domains
- LLM-generated noisy text: detection, characteristics, and implications
- Instruction-following in LLMs when faced with noisy or ambiguous prompts
- Retrieval-augmented generation (RAG) with noisy documents
Real-World LLM Applications
- LLM performance on text from social media, forums, and messaging platforms
- Handling code-switching, slang, and emerging language in LLMs
- LLMs for content moderation and safety in user-generated contexts
- Multilingual and cross-lingual LLM applications on informal text
- LLMs for assisting language learners and processing learner text
- Bias, fairness, and representation in LLMs trained on user-generated text
We particularly encourage submissions that address multilingual challenges, low-resource languages, cross-platform variations, and the unique characteristics of user-generated text across different communities.
Submissions should conform to the ACL style guidelines. Long and short paper submissions must be anonymized. Please submit your papers via:
Submit via OpenReview Commit via ARR
Double Submission Policy: Papers that have been or will be submitted to other meetings or publications must indicate at submission time. Authors of a paper accepted for presentation must notify the workshop organizers by the camera-ready deadline as to whether the paper will be presented or withdrawn.
EMNLP 2026 Findings:If you would like to present your EMNLP findings paper at WNUT, please fill out the following form by TBA.
Organizers
Contact
jy.bak@skku.eduProgram Committee
We thank the W-NUT 2026 program committee members for their contributions to the review process. Members are listed alphabetically by first name.
-
Aditya Jain
-
Advitya Gemawat
-
Alba Pérez-Montero
-
Albina Sarymsakova
-
Alexander Mehler
-
Alistair Plum
-
Anay Abhijit Dombe
-
Ankit Bhattacharjee
-
Antonios Anastasopoulos
-
Ashfaq Ali Shafin
-
Avinash Goutham Aluguvelly
-
Bitan Majumder
-
Bogdan Savelyev
-
Bolei Ma
-
Chao Jiang
-
Claire Wang
-
Cooper Ballard
-
Danae Sanchez Villegas
-
Danilo Croce
-
Danni Liu
-
Diana Inkpen
-
Eduardo Blanco
-
Emily Allaway
-
Eshaan Jain
-
Fayeq Jeelani Syed
-
Federico Ruggeri
-
Gianmarco Pappacoda
-
Günter Neumann
-
Hari Kang
-
HyunJin Kim
-
Iñaki San Vicente
-
Ishan Jindal
-
Jaehyeok Lee
-
Jane Arleth dela Cruz
-
Julia Mendelsohn
-
Julina Maharjan
-
Khandaker Mamun Ahmed
-
Krityapriya Bhaumik
-
Linsey C. Yang
-
Lucy H. Lin
-
Manuel Montes
-
Manuel R. Ciosici
-
Maria Antoniak
-
Mariët Theune
-
Marko Haralovic
-
Markus J. Hofmann
-
Mike Zhang
-
Mirco Schönfeld
-
Monica Hegde
-
Muhammad Mubashir Hassan
-
Naoki Otani
-
Nathanael Chambers
-
Nelu D. Radpour
-
Nikola Ljubešić
-
Nils Schwager
-
Noor Mairukh Khan Arnob
-
Onat Dogan Akca
-
Phuong-Anh Nguyen-Le
-
Piyush Agarwal
-
Raghvi Baloni
-
Rajeel Ahmad Ansari
-
Rajveer Singh Pall
-
Richard Sproat
-
Ruslana Margova
-
Sai P Vallurupalli
-
Sai Tulasi Kolapudi
-
Saman Rahbar
-
Sangmin Song
-
Seongwoo Choi
-
Shreyash Rawat
-
Sivaraaman Balakrishnan
-
Souvik Das
-
Suhyeon Park
-
Tanmoy Mazumder
-
Tanvir Ahmed Sijan
-
Ted Zhang
-
Vincent Ng
-
Xiliang Zhu
-
Xinyu Li
-
Xuanjing Chen
-
Yangfeng Ji
-
Yashdeep Thorat
-
Yasuhide Miura
-
YeongJun Hwang
-
Yuval Pinter
-
Zihao Zheng