Attention Is the New Big-O: A Systems Design Approach to Prompt Engineering (alexchesser.medium.com)

Hi HN, we are Zaid, Muhammad and Hammad, the co-founders of Uplift AI (https://upliftai.org). We build models that speak underserved languages — today: Urdu, Sindhi, and Balochi.

A billion people worldwide can't read. In countries like Pakistan – the 5th most populous country – 42% of adults are illiterate. This holds back the entire economy: patients can't read medical reports, parents can't help with homework, banks can't go fully digital, farmers can't research best practices, and people memorize smartphone app button sequences. Voice AI interfaces can fix all of this, and we think this will perhaps be one of the great benefits of modern AI.

Right now, existing voice models barely work for these languages, and big tech is moving slowly.

Uplift AI was originally a side project to make datasets for translation and voice models. For us it was a "cool side-thing" to work on, not an "important full-time thing" to work on. With some initial data we hacked together a Urdu Voice Bot on Whatsapp and gave it to one domestic worker. In two days 800 people were using it. When we dived deeper into understanding the users, we learned that text interfaces don't work for sooo many. So we started Uplift AI to solve this problem fulltime.

The most challenging part is that all the building blocks needed for great voice models are broken for these languages. For example, if you are creating a speech synthesis model, you will scrape a lot of data from youtube and auto-label it using a transcription model… all very easy to do in English. But it doesn't work in under-served languages because the transcription modes are not accurate.

There are many other challenges. Like when you hire human transcribers to label the data, often they don't have any spell correctors for their languages, and this creates lots of noise in the data… making it hard to train models with low data. There are many more challenges in phonemes, silence detection, diacritization etc.

We solve these problems by making great internal tooling to help with data labeling. Also, we source our own data and don't buy it. This is counterintuitive, but a big advantage over companies buying data and then training. By sourcing our own data we create the right data distributions and get much better models with much less data. By doing the entire thing inhouse, (data, labeling, training, deploying) we are able to make a lot faster progress.

Today we publicly offer a text to speech APIs for Urdu, Sindhi, and Balochi. Here's a video which shows this: https://www.loom.com/share/dcd5020967444c228e9c127151e7a9f5.

Khan Academy is using our tech to dub videos to Urdu (https://ur.khanacademy.org).

Our models excel at informational use cases (like AI bots) but need more work in emotive use-cases like poetry.

We have been giving a lot of people private access in beta mode, and today are launching our models publicly. We believe this will be the fastest way for us to learn about areas that are not performing well so we can fix them quickly.

We'd love to hear from all of you, especially around your experiences with under-served languages (not just the Pakistani ones we're starting with) and your comments in general.

Comments (17)

_waqas_ali_ · 5m ago

As a Sindhi speaker myself, amazing stuff. The output is so good. This unlocks the vastness of the internet for millions of people. I am imaging something like NotebookLM but for under-served languages or a hotline where people can call and talk/learn about anything. Do you guys have plans to create b2c products yourself?

pavlov · 1h ago

Nice! Clearly a big and underserved market for voice AI solutions.

Would be nice to have some code examples for using your TTS API with Pipecat.

zaidqureshi · 1h ago

I have to make that.. I did make one for LiveKit which utilizes our websocket API designed for real-time conversation API:

https://docs.upliftai.org/tutorials/livekit-voice-agent

zaidqureshi · 33m ago

btw I did try to first make it with Pipecat and was having some annoying windows issues with getting libraries installed for daily etc. so I posted something that was easily reproducible for the tutorial...

nojs · 42m ago

Nice, this is really needed. Would be cool to see some of the less common regional Chinese dialects, which are widely spoken and often the only language older people speak. And even just more accurate regional accents for Mandarin.

zaidqureshi · 38m ago

wow did not know that! Do you feel there is gap in speech understanding here or personalization missing with current TTS?

moinism · 47m ago

Congrats on the launch! Having support for regional voices is going to open up so many opportunities.

zaidqureshi · 39m ago

Agreed!

akshayp29 · 2h ago

Pretty cool! Do you think the model would be good at other under-served languages as well? Or is it hypertuned to just these?

zaidqureshi · 2h ago

The model itself can work well for new languages, its just the process of data gathering and maintaining high quality of data is what we have to figure out as we scale across languages.

Currently the model is only given data for these languages so it doesn't know anything else.

mandeepj · 24m ago

> just the process of data gathering and maintaining high quality of data is what we have to figure out as we scale across languages.

À crawler and days ingestion pipeline will not help with that?

zaidqureshi · 19m ago

Gathering audio data online is not that hard, but getting it accurately labelled is challenging, as the speech understanding systems for those languages aren't there either, so we can't automatically do that

akshayp29 · 1h ago

Cool - makes sense!

sanman8119 · 1h ago

Would love to see Malayalam here one day!

zaidqureshi · 1h ago

Yes! I will keep track of this comment for the day we do :P

yorwba · 1h ago

Unless that happens within a week or so, this thread will be locked and you won't be able to reply anymore.

It would be good to have a company blog with an RSS feed that people can subscribe to for updates.

zaidqureshi · 1h ago

ah, created a quick google form for language requests! https://forms.gle/XA6nZbmBNK5K7GJv5

Without the Futex, It's Futile (h4x0r.org)

Launch HN: Parachute (YC S25) – Guardrails for Clinical AI

Custom telescope mount using harmonic drives and ESP32 (svendewaerhert.com)

Critical Cache Poisoning Vulnerability in Dnsmasq (lists.thekelleys.org.uk)

Launch HN: Uplift (YC S25) – Voice models for under-served languages

Google President Praised MAGA Speech Slamming 'Climate Extremist Agenda' (desmog.com)

Lazy-brush – smooth drawing with mouse or finger (lazybrush.dulnan.net)

Chrome intends to remove XSLT from the HTML spec (github.com)

Prime Number Grid (susam.net)

Candle Flame Oscillations as a Clock (cpldcpu.com)

PyPI Preventing Domain Resurrection Attacks (blog.pypi.org)

OpenMower – An open source lawn mower (github.com)

Vim Macros for Beancount (tangled.sh)

Geotoy – Shadertoy for 3D Geometry (3d.ameo.design)

How to Build a Medieval Castle (archaeology.org)

In 2006, Hitachi developed a 0.15mm-sized RFID chip (hitachi.com)

EloqKV, a distributed database with Redis compatible API (GPLv2 and AGPLv3) (github.com)

Show HN: Whispering – Open-source, local-first dictation you can trust (github.com)

Ted Chiang: The Secret Third Thing (linch.substack.com)

Netflix Revamps Tudum's CQRS Architecture with Raw Hollow In-Memory Object Store (infoq.com)

Counter-Strike: A billion-dollar game built in a dorm room (nytimes.com)

Attention Is the New Big-O: A Systems Design Approach to Prompt Engineering (alexchesser.medium.com)

The Life and Death of London's Crystal Palace (2021) (heritagecalling.com)

Tiny-tpu: A minimal tensor processing unit (TPU), inspired by Google's TPU (github.com)

X-ray scans reveal Buddhist prayers inside tiny Tibetan scrolls (popsci.com)

Show HN: I built an app to block Shorts and Reels (scrollguard.app)

Left to Right Programming (graic.net)

Ask HN: Have any successful startups been made by 'vibe coding'?

Obsidian Bases (help.obsidian.md)

UK drops demand for backdoor into Apple encryption (theverge.com)

Spice Data (YC S19) Is Hiring a Product Associate (New Grad) (ycombinator.com)

How to Use Snprintf (bernsteinbear.com)

FFmpeg Assembly Language Lessons (github.com)

Why I'm all-in on Zen Browser (werd.io)

Intel 80286 emulator for Raspberry Pico (github.com)

Show HN: We started building an AI dev tool but it turned into a Sims-style game (youtube.com)

The rising returns to R&D: Ideas are not getting harder to find (papers.ssrn.com)

Croatian freediver held breath for 29 minutes (divernet.com)

An IRC-Enabled Lawn Mower (2021) (jotunheimr.idlerpg.net)

Anna's Archive: An Update from the Team (annas-archive.org)

Launch HN: Reality Defender (YC W22) – API for Deepfake and GenAI Detection (realitydefender.com)

Newgrounds: Flash Forward 2025 (newgrounds.com)

A Case for Protecting Computer Games with SGX (2016) (dl.acm.org)

Show HN: I built a toy TPU that can do inference and training on the XOR problem (tinytpu.com)

Typechecker Zoo (sdiehl.github.io)

Shamelessness as a strategy (2019) (nadia.xyz)

End well, this won't: UK commissioner suggests govt stops kids from using VPNs (theregister.com)

Association for the Preservation of Spiritualist and Occult Periodicals (iapsop.com)

The lottery ticket hypothesis: why neural networks work (nearlyright.com)

MCP doesn't need tools, it needs code (lucumr.pocoo.org)

Launch HN: Uplift (YC S25) – Voice models for under-served languages

Comments (17)