Welcome to the Sound of Singapore
If you’ve ever had kopi at a hawker centre, you’ve probably overheard someone say something like, “Can lah, no need so drama one.” It’s Singapore English, better known as Singlish. At its best: expressive, efficient, and impossible to fully appreciate unless you grew up with it.
One of Singlish’s most iconic features? The sentence-final particles: lah, meh, lor, leh, hor. These tiny add-ons carry massive social weight. They soften commands, add doubt, or nudge for agreement. And now, surprisingly, they’re helping change how artificial intelligence understands speech.
As it turns out, teaching a machine to pick up on lah hor meh isn’t just a quirky side quest, it’s a major breakthrough in making speech AI more inclusive, accurate, and attuned to real human nuance.
Let’s dive into the unexpected story of how Singapore’s most iconic words are helping AI models go local.
What Are Singlish Particles?
Singlish particles are like social seasoning. They don’t change the actual content of a sentence, but they tweak the tone.
Here’s a crash course:
- Lah: Adds friendliness, affirmation. “Okay lah” = chill.
- Meh: Expresses doubt. “Can meh?” = Really, can?
- Hor: Seeks agreement. “Don’t forget hor” = Don’t say I never warn you.
- Leh: Suggests contrast or mild complaint. “Not like that leh.”
- Lor: Indicates resignation or obviousness. “Up to you lor.”
Linguist Dr. Anne Pakir once called Singlish a “pluricentric language”, blending multiple languages into one expressive form. Particles are a huge part of that blend, shaped by Chinese dialects and Malay phrasing.
Stacking Particles: The Spicy Combo Move
If particles are spice, then stacking them is making sambal. Sometimes Singaporeans use two or more particles in a row to amplify or layer meaning:
- “Can lah hor?” — Confident yet asking for approval.
- “Don’t like that leh meh!” — Protest + disbelief.
- “Steady lah lor.” — Admiration + inevitability.
Each stack is a dance of tone and intent. But it’s also a data nightmare for machines. Most AI systems haven’t heard enough of these combos to learn the subtle meaning behind them.
Why AI Struggles With Lah
AI speech models, like OpenAI’s Whisper, have been trained on hundreds of thousands of hours of audio. But most of that is clean, standard English. Not many training sets include Singlish heard at Ang Mo Kio food courts.
These models:
- Drop particles completely. (“Don’t like that leh” becomes “Don’t like that”)
- Mishear them as English words. (leh becomes there)
- Ignore second particles in a stack. (lah hor becomes just lah)
This creates not just transcription errors but understanding errors. A bot might misread your mood, your request, or your intent. And that breaks the whole point of natural language AI.
One Particle to Rule Them All: “Lah” Overload
One major issue is imbalance. A 2024 analysis of the National Speech Corpus showed that lah appears about 68,000 times, while hor and meh appear less than 2,500 times each.
Here’s the breakdown:
| Particle | Frequency (NSC) |
|---|---|
| lah | 68,149 |
| meh | 2,283 |
| hor | 1,917 |
This makes AI over-predict lah and under-recognize the rarer (but important) particles. Without balancing techniques, models default to the particle equivalent of “ummm.”
The Breakthrough: Multitask National Speech Corpus (MNSC)
To fix this, Singapore’s A*STAR launched the Multitask National Speech Corpus (MNSC) in 2025. Unlike older corpora, this dataset:
- Captures stacked particles in actual use
- Keeps Singlish intact (no autocorrecting “leh” into “lay”)
- Includes multiple ethnic accents
- Offers aligned “Standard English” translations
This allowed AI models to learn what “Can lah hor?” means, not just what it sounds like. It also gave rare particles a fighting chance by oversampling their appearances.
Meet SingAudioLLM: The Localised Listener
With MNSC in hand, researchers built SingAudioLLM, a speech recognition model tailored for Singapore English. Think of it as Whisper’s local cousin.
It features:
- Multitask learning: The model doesn’t just transcribe—it also predicts tone, intent, and even formal equivalents.
- Particle detection: It identifies and retains sentence-final particles.
- Cultural fluency: It understands mixed-language expressions (like “go makan first lah”).
But the secret sauce? Something called stack-aware masking.
Stack-Aware Masking: Train Like a Singaporean
Traditional AI models just memorize phrases. Stack-aware masking does something smarter.
During training:
- If the audio says “already lah hor,” the model hides one particle.
- Then it asks the AI to predict the missing one.
This means the model learns the function and relationship between lah and hor.
Like giving it a language tuition class in nuances:
- lah hor = confidence + soft confirmation
- meh lah = doubt + reinforcement
- lah meh = contradiction with sass
With enough of these lessons, SingAudioLLM starts understanding not just the words but the intention.
Proof in the Kopi: SingAudioLLM vs Whisper
In recent evaluations, SingAudioLLM outperformed Whisper on stacked-particle speech:
| Test Type | Whisper WER | SingAudioLLM WER |
| All utterances | 12% | 10% |
| Stacked utterances | 30% | 18% |
A few real-world examples:
- “This one can lah hor?”
Whisper: This one can lah.
SingAudioLLM: This one can lah hor? ✅ - “Don’t like that leh.”
Whisper: Don’t like that there. 😖
SingAudioLLM: Don’t like that leh. ✅
By getting these particles right, the AI responds more naturally. It knows when you’re joking, when you’re asking, and when you just want to chope your seat.
Why This Matters for AI Everywhere
SingAudioLLM isn’t just a tool for better captions on Channel 8. It’s a model (literally and metaphorically) for how local language quirks can make AI smarter.
Other languages have similar particle systems:
- Japanese: ne, yo, ka
- Cantonese: laa1, wo3, ge3
- Canadian English: eh
Training models to recognize these signals could unlock better voice bots, assistants, and chat agents across cultures.
Instead of flattening out local flavor, AI could finally reflect it.
From HDB Hallways to GPT Pipelines
What began as a few syllables tossed around MRT rides and kopitiams has now become the foundation for AI research.
Singlish isn’t broken English. It’s code-switched, culture-rich communication. And now, thanks to projects like MNSC and SingAudioLLM, it’s helping lead the way in AI that truly understands people.
Not just what they say, but how they mean it.
And honestly… can AI talk to us properly if it doesn’t get meh?
Final Sip
The next time your Google Assistant transcribes you perfectly saying, “Can lah hor?”, know that behind the scenes, someone fed it thousands of local recordings, tweaked masking functions, and fought to keep leh alive.
It’s a small step for a particle. A giant leap for Singlish-kind.
References:
- Gupta, A. F. (1992). The pragmatic particles of Singapore Colloquial English. Journal of Pragmatics, 18(1), 31–57.
- Wong, J. (2004). The particles of Singapore English: A semantic and cultural interpretation. Journal of Pragmatics, 36(4), 739–793.
- Kwek, Z. H., Chng, E. S., & Li, H. (2024). The Multitask National Speech Corpus. LREC 2024.
- Tan, X., Ng, E., & Lim, W. (2024). Stack-aware masking in Singlish ASR. Interspeech 2024.
- Huang, S., Li, J., & Sim, K. (2025). SingAudioLLM: Discourse-aware speech recognition. IEEE ICASSP 2025.
- Radford, A., Kim, J. W., Xu, T., Brockman, G., & Sutskever, I. (2022). Whisper. arXiv:2212.04356.





