Uzbek Founder Builds AI Translator for Turkic Languages, Startup Valued at $1.5M

Mukhammadsaid Mamasaidov was born in Tashkent and grew up in a family of IT professionals. He graduated from the Tashkent branch of South Korea’s Inha University with a degree in Software Engineering. In his second year at university, he became interested in natural language processing and built an autocorrect tool for Uzbek. That project later grew into Tilmoch, an AI-powered translator for Turkic languages.

The platform now has 200,000 monthly users. Since its launch, the startup has raised more than $300,000 in grants and venture funding, and in 2025 Tilmoch was valued at $1.5 million.

For the joint Digital Business and Astana Hub project “100 Startup Stories from Central Eurasia”, Mukhammadsaid spoke about why general-purpose LLM chatbots still struggle with Turkic languages, the challenges Tilmoch faced when entering the Kazakh market, and how the service plans to grow into a full-fledged super app.

Read this article in Kazakh

“We built a corpus of 35,000 Uzbek-language texts”

— Mukhammadsaid, what got you interested in AI translation and natural language processing? How did the project get started?

— In my second year at university, we had a course called Data Structures and Algorithms, and we had to build a project for it. At the time, I was frustrated that most operating systems didn’t natively support autocorrect for Uzbek, so I decided to create a tool that could automatically correct mistakes.

Uzbek is a complex language: lots of suffixes and other affixes can be added to a root, changing the form of the word. On top of that, Uzbek uses two alphabets, Latin and Cyrillic. People also speak Russian and English, and there are many different dialects. A lot of people use non-standard forms or make mistakes with words borrowed from Arabic. All of this makes it harder to digitize the language. It becomes difficult to use Uzbek effectively for online search or to build digital services, because people can write the same word in different ways. I wanted to bring these different spellings into a single standard.

— What did the first version of the autocorrect tool look like?

— It was very simple: just code, with no user interface at all. We took a list of Uzbek words and a list of around 150–200 possible endings, then combined them mechanically.

We checked the results against a dictionary: if a word form appeared in the dictionary or in the uploaded text, the algorithm treated it as correct. If it didn’t, it flagged it as an error. It was a pretty crude and incomplete approach. Users found it frustrating when a perfectly correct word was marked as wrong simply because that specific form wasn’t in the database. Once the project was over, I wanted to build a proper system that knew all the rules of Uzbek and could understand the structure of each word algorithmically.

— How did a project focused on Uzbek autocorrect turn into a startup?

— Back then, I had no idea how startups worked. At first, we wanted to build a kind of Grammarly for Uzbek: a website where people could paste in their text and check it for mistakes. I built the first web version, but no one used it. It became clear pretty quickly that writing code wasn’t enough. We also needed to focus on marketing and promotion.

Then I heard about mGovAward, a mobile app competition organized by IT Park Uzbekistan together with the Ministry of Digital Technologies and the UAE government. To enter, I needed a team, so I brought in two of my classmates, a designer and a data specialist. In spring 2022, we developed a mobile keyboard for Uzbek with predictive text and autocorrect. We didn’t have a linguist on the team, so I had to work through grammar textbooks myself and figure out how 150–200 different types of affixes work and how they change the root of a word. We also drew on research into Turkish and Finnish and consulted specialists. Using dictionaries, grammar rules and analysis of news texts, we trained and refined the model step by step. The project ended up taking first place and won us $50,000.

— What happened to the project after you won the competition?

— We soon realized that a mobile keyboard was extremely difficult to monetize, so we decided to build a more advanced Uzbek spellchecker called Tahrirchi. The idea was to make it a Microsoft Word plugin. Word is used by almost every organization in Uzbekistan, and for most people it is much more convenient to check documents while they are working on them.

At the same time, we started working on more complex grammar and style corrections, as well as punctuation. To train the model, we needed a large digital corpus of Uzbek, but there simply wasn’t enough data available at the time. Relying only on news websites wasn’t an option because the language there was too repetitive and the vocabulary too limited. We realized that the data we needed was sitting in printed books in Uzbek.

We used whatever digital editions were available, while rarer books had to be tracked down in shops, bought from second-hand booksellers and scanned. In the end, we managed to assemble around 35,000 digitized texts into a single language corpus. We then took the product to the public sector, but working with government institutions turned out to be difficult, so we kept looking for the right product-market fit.

At the end of 2023, we entered the President Tech Award competition. The project won in the Artificial Intelligence category and received $100,000. That funding gave us the chance to rethink the product completely and find a niche that people were actually willing to pay for. We looked at speech recognition, voice bots and computer vision solutions. But when we analyzed Google search data, we saw huge demand for high-quality translation. At the time, ChatGPT and Claude were still weak in Uzbek, while Google Translate often produced unnatural, machine-like text. We spent all of 2024 developing a specialized translation tool, using the grant money for servers, model training and marketing. Tilmoch was publicly launched in September 2024.

— How is Tilmoch’s translation different from general-purpose models like ChatGPT or Claude?

— Large language models fundamentally operate in English. Turkic languages are low-resource languages, which means there is very little digitized material available for AI models to train on. As a result, text generated in Uzbek or Kazakh often sounds unnatural, with English sentence structures coming through. General-purpose chatbots can also miss parts of the text when translating long documents, distort the meaning or stop generating halfway through a sentence.

Tilmoch is built specifically for translation, so it doesn’t drop parts of the text or distort the meaning. We regularly run blind tests with professional linguists, comparing the results from our model with Google Translate and LLM chatbots. In 80–90% of cases, the experts prefer Tilmoch’s translations because the wording sounds more natural and the phrasing is more precise.

To get this level of quality without spending huge amounts on servers, we built two models: a lighter, faster one and a heavier, more advanced one. The basic model translates text instantly and handles everyday tasks. If users want to improve the translation, there are dedicated tools in the interface. The “Polish” button activates the heavier neural network, which looks at the context of the entire paragraph, checks it against our translation memory and finds more natural ways to phrase it. If the original translation is already accurate, the algorithm leaves it as it is. If it spots an awkward phrase, it rewords it.

“After San Francisco, our team became much more ambitious”

— In September 2025, you raised $150,000 from three venture funds. How did you find the investors?

— After the official launch, Tilmoch gradually started gaining traction with different groups of users. The service became popular among civil servants, lawyers, marketers, bloggers and SMM specialists. Revenue kept growing steadily, and by the time we started talking to venture funds, monthly recurring revenue had reached $8,000–9,000.

We caught the attention of Davron Parmanov, an operating partner at the venture fund Aloqa Ventures. He got in touch and suggested raising an investment round. Aloqa Ventures became the lead investor and brought in two other local funds, IT Park Ventures and Yoshlar Ventures. The startup was ultimately valued at $1.5 million. Each of the three funds invested between $40,000 and $60,000, bringing the total round to $150,000.

— In 2025, you also took part in the Silicon Valley Residency and the AlchemistX accelerator in the US. What did you get out of that experience?

— The program was run with the participation of Astana Hub, IT Park Uzbekistan, Alchemist Accelerator and Silkroad Innovation Hub. Astana Hub Ventures and IT Park Ventures received minority stakes in the startup in exchange for access to the program, mentorship and international opportunities.

Three months in San Francisco changed the way we looked at the business. The AlchemistX accelerator focused mainly on B2B sales. Every week, experienced founders ran workshops, we worked with a dedicated Alchemist mentor, and the program ended with a classic Demo Day where we pitched our projects to US investors.

The biggest takeaway I brought back from Silicon Valley was a completely different sense of scale and attitude to risk. In Central Asia, $1 million in revenue is often seen as the ultimate goal. In the US, it is just the starting point. There is also an incredible concentration of talent. On industry platforms, you can find around a dozen events happening every evening where you can meet people, make connections and start building projects together. US investors are willing to back even very early-stage ideas if the founder is confident and genuinely passionate about what they are building. The trip raised the team’s ambitions considerably.

— How does the monetization model work? How much revenue does the startup generate, and which segment brings in the most?

— Our current MRR is $22,000–23,000, and the goal is to reach $30,000 by the end of the year. Most of our revenue comes from the B2C segment.

On the B2C side, we have more than 200,000 monthly active users in Uzbekistan and around 1.2 million visits. About 98% of the audience uses the platform for free, with a limit on the number of characters. We keep that traffic intentionally because it gives us a steady flow of data to further train the models and helps us expand our reach. Around 4,000–4,200 users pay for a subscription on a regular basis. The basic plan costs $4 per month, while the top-tier plan is $15.

In the B2B segment, we work with around 150 companies in Uzbekistan. Corporate clients buy team plans starting at $200 per user per year.

Our B2B clients include the localization team at Yandex Go, as well as streaming services Kinopoisk and iTV, which use the product to translate subtitles. We also have an AI voice-over technology. It is not yet suitable for dubbing feature films or TV series, but one of our clients is already using it successfully to translate educational courses into Uzbek.

For more complex integrations, API access costs $500–1,000 per year. For larger businesses, we develop individual gateways and custom solutions tailored to their needs. We also prepare labeled datasets for banks that are training their own internal language models in Uzbek.

“Turkic languages have a critical shortage of ready-made text corpora”

— How did you expand into Kazakh and other languages?

— We thought we had a universal formula that worked. For each new language, we did alignment, matching parallel texts sentence by sentence. In practice, that meant taking the same book in Russian, English and the target language and lining up the sentences so the model could see how the same idea was expressed in different languages. In 2025, we used this approach to add Kazakh, training on around 4,000 books and texts from open sources. But the translation quality turned out to be fairly weak.

To understand what was going wrong, we did some customer research with professional linguists, translators and copywriters from Kazakhstan. It turned out that a large share of Kazakh-language content online and in modern printed books contains machine-translated text and phrasing copied too closely from Russian. The model was learning from that data, so our translations ended up sounding AI-generated too. We ran into a similar problem with other Turkic languages.

Expanding into Kazakhstan was also difficult because we weren’t physically based there, which made it harder to understand how users were responding to the product. We have since fixed the issues with the model, and the translations now sound more natural. But for the Kazakh market, we are considering entering through related tools such as dictionaries or copywriting review services. At the moment, we have around 20,000 monthly users in Kazakhstan, with about 300 paying for a subscription.

— Tilmoch currently supports 16 languages, including Azerbaijani, Georgian and even Korean. Are you planning to add more languages, and what are your goals for the next year?

— Adding a new language from scratch and getting the quality right takes three to four months. For now, we have decided to pause expansion and focus all our efforts on improving translation quality for Kazakh, Kyrgyz, Tajik, Karakalpak and Azerbaijani. There is still a critical shortage of ready-made text corpora, so we are now using data augmentation. We generate the syntactic structures we need with third-party LLMs, check the quality, and then use that data to further train our own model.

Our broader goal is to turn the service into a language super app. We have already added word puzzles that more than a thousand people play every day, and launched audio transcription in Uzbek and Kazakh. The system now processes 50–60 hours of audio per day. We are also developing a portal with spelling rules and tools for generating standard documents and contracts.

Over the next year, we plan to launch mobile keyboards for the new Latin-based Uzbek alphabet to help people get used to it more quickly, further digitize the language and make it easier to use.

Our strategic goal is to grow traffic from the current 1.2 million to 5 million visits per month. Once we build an audience of several million users and strong, high-quality traffic, we may consider a potential deal or integration with one of the region’s major digital ecosystems.