About
I'm a third year PhD student at the Center for Language and Speech Processing at Johns Hopkins University, advised by Professor David Yarowsky. In the past, I worked at ALMAnaCH, INRIA in Paris as a Research Engineer with Benoît Sagot and Rachel Bawden. Even before that, I graduated from the EMLCT Masters' program as an Erasmus scholar, with a dual MSc. in Computational Linguistics and NLP at Charles University, Prague (first year) and Saarland University, Saarbrücken (second year). I'm interested in building NLP tools for text and speech that are available for all the world's languages in their dialectal, colloquial, and code-switched variants :)
Publications
Here's a fun visualization of some of my work!
Or see ACL Anthology or Google Scholar or this list.
-
There is No Theoretical Curse of Multilinguality For Embedding Space Structure .
Niyati Bafna, Neha Verma, Vilém Zouhar, Philipp Koehn, David Yarowsky.
-
Multilingual Reasoning Cascades Need More Context .
Arnav Mazumder, Dengjia Zhang, Shuyue Stella Li, Yulia Tsvetkov, Niyati Bafna. -
Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning .
Niyati Bafna, Ryan Soh-Eun Shim, Barbara Plank, David Yarowsky, and Hale Sirin.
EMNLP 2026 (main). -
Omnilingual MT: Machine Translation for 1,600 Languages .
Belen Alastruey*, Niyati Bafna*, Andrea Caciolai*, Kevin Heffernan*, Artyom Kozhevnikov*, Christophe Ropers*, Eduardo Sánchez*, Charles-Eric Saint-James*, Ioannis Tsiamas*, Chierh Cheng, Joe Chuang, Paul-Ambroise Duquenne, Mark Duppenthaler, Nate Ekberg, Cynthia Gao, Pere Lluís Huguet Cabot, João Maria Janeiro, Jean Maillard, Gabriel Mejia Gonzalez, Holger Schwenk, Edan Toledo, Arina Turkatenko, Albert Ventayol-Boada, Rashel Moritz, Alexandre Mourachko, Surya Parimi, Mary Williamson, Shireen Yates, David Dale, and Marta R. Costa-jussà.
EMNLP 2026 (main). -
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models .
Emily Chang and Niyati Bafna.
ACL 2026 (main). Oral presentation. Considered for award. -
How Important is `Perfect' English for Machine Translation Prompts? .
Patrícia Schmidtová*, Niyati Bafna*, Seth Aycock*, Gianluca Vico, Wiktor Kamzela, Katharina Hämmerl, Vilém Zouhar.
EACL Findings 2026. -
The Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation Failure .
Niyati Bafna, Tianjian Li, Kenton Murray, David R. Mortensen, David Yarowsky, Hale Sirin, Daniel Khashabi.
AACL 2025 (main). -
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
Niyati Bafna and Matthew Wiesner.
Interspeech 2025. -
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models.
Niyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen, Kenton Murray, David Yarowsky, and Hale Sirin.
ACL 2025 (main). -
Evaluating Large Language Models along Dimensions of Language Variation: A Systematik Invesdigatiom uv Cross-lingual Generalization.
Niyati Bafna, Kenton Murray, and David Yarowsky.
EMNLP 2024 (main). -
Pointer-Generator Networks for Low-Resource Machine Translation: Don’t Copy That!
Niyati Bafna, Philipp Koehn, and David Yarowsky.
Fifth Workshop on Insights from Negative Results in NLP, co-located with NAACL 2024. When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages.
Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, and Rachel Bawden.
COLING 2024. Best Student Paper Award.Cross-Lingual Strategies for Low-Resource Language Modeling: A Study on Five Indic Dialects.
Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, and Rachel Bawden.
TALN 2023Combining Noisy Semantic Signals with Orthographic Cues: Cognate Induction for the Indic Dialect Continuum.
Niyati Bafna, Josef van Genabith, Cristina España-Bonet, and Zdeněk Žabokrtský.
Conference on Computational Natural Language Learning (CoNLL) 2022.Subword-based Cross-lingual Transfer of Embeddings from Hindi to Marathi and Nepali.
Niyati Bafna and Zdeněk Žabokrtský.
SIGMORPHON 2022, co-located with NAACL 2022.Clause Final Verb Prediction in Hindi: Evidence for Noisy Channel Model of Communication.
Kartik Sharma*, Niyati Bafna*, and Samar Husain.
Workshop on Cognitive Modeling and Computational Linguistics (CMCL) 2021.Towards Universal Segmentations: UniSegments 1.0.
Zdeněk Žabokrtský, Niyati Bafna, Jan Bodnár, Lukáš Kyjánek, Emil Svoboda, Magda Ševčíková, and Jonáš Vidra.
LREC 2022.Constrained Decoding for Technical Term Retention in English-Hindi MT.
Niyati Bafna, Martin Vastlik, and Ondrej Bojar.
International Conference on Natural Language Processing (ICON) 2021.Towards Handling Verb Phrase Ellipsis in English-Hindi Machine Translation.
Niyati Bafna and Dipti Sharma.
International Conference on Natural Language Processing (ICON) 2019.
News
July 2026 We're building the "Last Translation Benchmark". Contribute your favourite hard-to-translate examples here.
April 2026 I passed my GBO. Here's my proposal.
May 2025 I'm interning at Meta FAIR (again) to work on Omnilingual MT for 1600 languages.
May 2025 Check out dialup, a Python package that lets you do fun things like artificial dialect generation and making your MT systems robust to dialectal variation! Read the paper for more details.
April 2025 Gave an invited talk at Schmidt Sciences AI and Advanced Computing: Dialectal robustness in Large Language Models.
May 2024 Our work on Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages won the COLING Best Student Paper Award at LREC-COLING.
Feb 2024 I'll be interning at Seamless Meta in Menlo Park over the summer.
April 2023 Accepted my PhD offer at JHU CSLP, and will be starting in Fall 2023. I'll be advised by Professor David Yarowsky.
Oct 2022 Invited talk at Linguistic Mondays, Institute of Formal and Applied Linguistics, Charles University, on two of my recent works in experiments on Indic languages: Empirical Models for an Indic Dialect Continuum
Oct 2022 Starting as a research engineer at ALMAnaCH, INRIA in Paris, with Benoît Sagot and Rachel Bawden; super excited :)
Aug 2022 Defended my Masters thesis (twice), graduated from Charles University and Saarland University. I did my thesis jointly with the MLT group at DFKI and UFAL. I was supervised by Prof. Josef van Genabith and Cristina España-Bonet from the former and Zdeněk Žabokrtský from the latter. The thesis is about cognate induction and data collection for 26 (extremely) low resourced languages of the Indic dialect continuum; check it out here: Empirical Models for an Indic Language Continuum!
Dec 2021 New paper at ICON '21, NIT Silchar, about constrained deconding for technical terms in English-Hindi MT with UFAL.
Apr 2022 New paper at CMCL@NAACL '21 with Prof. Samar Husain at IIT Delhi, India, about computational modelling of cognitive hypotheses, specifically, the adaptability hypothesis and noisy channel hypothesis.
Oct 2020 Started at the EMLCT Masters' program with an Erasmus scholarship.
May 2020 Graduated with a bachelors' degree from Ashoka University, and the Department of English prize :-)
Dec 2019 New paper at ICON 19, IIIT Hyderabad, with Prof. Dipti Misra Sharma at LTRC, about verb phrase ellipsis handling in English-Hindi MT.
CV
Here's a PDF version of all of this stuff.
Research Interests
There are 3800+ written languages in the world, with varying levels of resourcedness. Given the LLM paradigm that powers everything these days, making NLP massively multilingual has two broad facets: enabling LLMs to comprehend content and instructions in a low-resource language, and teaching it to generate accurate, useful, and fluent content in a low-resource language (LRL). Here are some of the kinds of problems in this space:
Answering the (age-old) cascade question: At the frontier of NLP today, we're able to do quasi-magical things in a few high-resource languages, especially English. For the rest of the languages in the world, our tools lag behind: they are more likely to produce wrong-language text, and their responses are less accurate. Given complex problems, they are less likely to make it through a series of logical steps correctly. They are less easily controlled for toxicity and safety. What we do have for these languages is good quality machine translation (well, better quality). Instead of attempting to induce native capabilities for all the above in LLMs for the range of mid-resource languages - why not translate inputs into English, let LLMs do what they know best, and then translate English outputs back into the target language? What do we lose by this cascaded approach, and what do we gain, and can we quantify these gains and losses? What does it mean for the kinds of resources we try to collect in LRLs, and for the tools we try to build for them? Should we redirect our energy towards building MT specialised in LLM outputs, instead of training LLMs in various languages?
Reasoning and inference-time scaling in a multilingual context: Once our LLM is pretained, we have a number of ways of making it better at complex tasks. A desired answer may not be straightforwardly derivable from the question for multiple reasons: the task itself might be difficult to understand, it may require a few logical hops, it may require a non-trivial specialised computation, or it may require consideration of several perspectives. Strategies such as few-shot prompting with in-context learning, chain-of-thought elicitation (and its multiple variants), tool use, and multi-agent debate, are methods of searching a solution space for a desired answer in response to the above problems. However, they are all worse in multilingual settings. (This by itself is somewhat interesting: if LLMs are doing all their reasoning in English, as some papers indicate they are, this shouldn't be the case.) I'm interested in gaining insights on this phenomenon of degradation - is it best characterised as coming from divergent base-model processes, shallower fluency and articulation problems, lack of post-training exposure, or something else? - and I'm interested in developing solutions to mitigate it.
Tool use for multilinguality: LLMs are still often bad at comprehension and generation for several mid- to low-resource languages. For these languages, we may want to leverage the reasoning capabilities of LLMs but outsource linguistic processing and generation to specialized modules. What are the best architectural and training strategies for these modules and their integration with the LLM pipeline?
Better machine translation: Yep, MT is not solved. MT is a moving target even for mid-to-high resource languages, because our horizons are constantly expanding with what we expect from NLP. Today we want to do fine-grained arguments, long coherent documents, literary generation, abstract correction. Even more audaciously: we want to do MT for *all* languages - we want to take lexicons and grammar books and show them to LLMs, and ask LLMs to translate Kalamang for us. How do we evolve MT for mid-resource languages to adapt it to new horizons? How do we teach LLMs to translate into and from unseen languages? How do we make all this robust and reliable? Further, MT evaluation very quickly becomes a problem: how do you holistically evaluate MT for a language you can only do string matching in? I'm interested in exploring MT catered to the frontier of NLP, for languages of various resource levels.
(And of course) Dialectal generalisation, cross-lingual transfer and its limits: Many dialects and languages of the world exist across a continuum, with varying degrees of resourcedness at various points. Can we model this continuum in a manner that can help NLP tools fill in the gaps"? How can we most effectively use its properties to leverage good datasets and models at certain points on this continuum for others? Many things about linguistic divergence are systematic, meaning that models can extrapolate performance to low-resource languages that are close enough to high-resource languages in regular ways. However, languages also have irregular, language-specific phenomena. Can we theoretically quantify the limits of cross-lingual transfer for a given language family? Can we identify the kinds of phenomena that cannot be transferred, so that we can evolve targeted solutions for them?
I also have a particular interest in NLP for Indian languages! Also - besides all this stuff, people seem to be talking a fair bit about agentic stuff these days. I should figure out what's up with that.For Fun...
When I'm not working, I enjoy playing tennis, dancing WCS, salsa, and bachata, solving cryptic crosswords, reading, and writing. I also enjoy the occasional game of rapid chess, and am a fan of the Sicilian Dragon.
Other fun things about me: I played basketball for my city as a teenager, my FIDE classical chess rating is 1609, and I spent the summer of 2019 translating a book from Hindi to English. And I have a 1-year+ Spanish streak on Duolingo, and speak no Spanish.