Bahasa AppIndonesian practice that grades grammar and tone separately
Bahasa App is a daily writing trainer for someone who reads Indonesian well but still sounds translated. We designed and built it to grade every reply twice, once for grammar and once for fit with the person it is written to, and never to merge the two.

- Client
- Own product
- Sector
- Language learning
- Year
- 2026
- Platforms
- Web app, Mobile web
- Services
- Product and UX design, Web development
Short daily practice in writing Indonesian at the right level of formality. Each exercise puts the learner in a real situation, such as a WhatsApp message to a lecturer, a reply to a paying client or a joke in a friends' group chat, and asks for the reply that situation calls for.
Built for
- A learner who already understands written Indonesian and wants to sound like a person writing it.
- Someone who writes to lecturers, clients, family and friends in Indonesian and needs a different register for each.
The problem
Indonesian changes with the person you are writing to. A message to a lecturer, a client and a close friend uses different words for I, you, not and already, and mixing levels in one message gives a translation away at once.
A learner whose Indonesian already gets the message across feels no pressure to improve it, and a grammar checker adds none: a reply that is grammatical but wrong for the person passes as correct.
The app asks the learner to find a problem before seeing the fix. If a correction reaches the browser early, it is one developer tools panel away, and the rule becomes an honour system.
Goals
- 01Grade grammar and register fit as two separate results, never one score.
- 02Name the mistake the app exists for: grammatical, but wrong for the person.
- 03Ask questions on the first attempt and show corrections only on the second.
- 04Grade what code can decide at once, offline, and queue the rest without guessing.
- 05Report progress as counts with their totals, with no streaks, points or overall score.
Our role
We designed and built the app, from the research behind its teaching rules to the interface.
What we delivered
- Research documents on teaching method, session flow and architecture
- React and Vite front end with light and dark themes
- Hono API with Drizzle and PostgreSQL
- Offline checker for register clashes, loanword spelling, English-shaped sentences and verb prefixes
- Practice scheduler that works on patterns, not single exercises
- Grading queue with a validated import for what the checker cannot decide
- Progress page, phrase bank and compose-time tracking
- Seed content: 8 patterns, 6 situations and 60 exercises
- Unit, API integration and Playwright end-to-end tests
How we worked
- 01
Research first
Three research documents on teaching method, session flow and architecture were written before the build, and the rules in the code name the document they come from.
- 02
Take the rules from a checked source
The formality vocabulary is transcribed from an Indonesian knowledge base. The table of standard loanword spellings is generated from 92 pairs checked against KBBI VI, the official dictionary, and is never edited by hand.
- 03
Enforce the rules below the interface
The two-attempt rule is held by one server function that builds every attempt response and by database constraints, and a probe cannot be stored unless it is phrased as a question.
- 04
Test the whole loop
Unit tests cover the checker and the scheduler, integration tests cover the API, and Playwright runs a session, a stilted verdict, the phrase bank, the progress page, an API outage and phone layouts.
Screens and features
Questions first, answers later
After a first attempt the learner gets questions about the problem spans, with no corrections, then writes again.

Progress on two separate axes
Grammar and register fit are counted separately, and the phrase bank keeps the same phrase at each level of formality.

Two readings, never averaged
Every reply gets a grammar result and a register result, each ok, minor or wrong. There is no combined score anywhere in the app.
Stilted, by name
A reply that is grammatical but wrong for the person is marked stilted. It leads the progress page, because it is the mistake the app exists to catch.
Questions first, answers later
The first attempt returns questions that point at the problem spans. The verdict, the corrections and a native model answer arrive only with the second attempt.
Situations set the difficulty
Six situations, from a message to a paying client to a joke in a close friends' group chat, decide the level of formality. Difficulty comes from who you are writing to, and the vocabulary stays ordinary.
Practice in short blocks
Each block drills one pattern with three to five fresh exercises. A pattern returns on a ladder of gaps from the same day to 45 days, depending on how its blocks go, and no prompt is ever shown twice.
Progress with its totals
Figures are counts with their totals, such as 4 of 10, next to compose time and error classes. There are no streaks or points, and pasted answers are left out of every timing figure.
A phrase bank with levels
Phrases are saved with their formality level and their source. The same phrase at a different level is a separate entry, because that contrast is the lesson.
Technical choices
What the product is built with, and why.
- A deterministic checker
- Register clashes, loanword spelling, English-shaped sentences and meN- prefix rules are checked in code in about a millisecond, with no model and no network. It covers register better than grammar, and register is the axis the learner needs most.
- A queue for what code cannot decide
- Whatever the checker cannot settle is written to a grading queue in PostgreSQL in the same transaction as the attempt. It is exported as files for a separate review step, and returning grades are validated against a schema, one transaction per file, so a bad file fails on its own. Until then the verdict says what was not decided offline.
- The answer withheld on the server
- One serializer builds every attempt response and leaves the corrections, the model answer and the verdict out of a first attempt. A database constraint refuses to record an answer as revealed before attempt two.
- A scheduler for patterns
- Practice is massed within a session and spaced across sessions, and a failed block moves its pattern back a step. The scheduler is pure functions with no database access, which is where its tests run.
- Compose time measured, never rewarded
- The composer records time to first keystroke and total compose time, and flags a paste so the server can leave it out of the medians. The scheduler never reads these times, so speed is reported and never rewarded.
- No model calls and no API key
- The app runs on local code and a local database. Nothing in the request path calls a model or needs a credential, so a practice session has no per-use cost.
Teaching something with no single right answer?
Grading on separate axes, and knowing what code can judge and what it cannot, are problems we have already worked through. Tell us what your learners practise.