← All work

Capsule

An asynchronous voice messenger with end-to-end encryption. You speak; they read one line.

Year
2026
Status
beta
Role
Solo developer — Android client, server, protocol
Stack
Kotlin · Rust ↔ Kotlin (uniffi) · MLS (RFC 9420) · SQLCipher · Go · PostgreSQL · protobuf · WebSocket · On-device ML
Goal
Voice messaging without pressure: the sender talks as long as needed, the recipient sees the gist at once and listens later.
Status
Fully working, in closed beta with friends.
Built
Android client, Go + PostgreSQL server, and a written protocol specification.

The idea

Voice messages are easy to record and tedious to listen to. Capsule transcribes speech and extracts the gist on the phone. The recipient reads one line and plays the audio later if they need to.

The path of a voice message
  1. VoiceA 0:42 recording
  2. TranscriptionSpeech to text on the device
  3. GistOne line, also on the device
  4. EncryptionBefore sending; only members hold the keys
  5. ServerStores and delivers, but cannot read
  6. RecipientDecrypts and reads one line

The server must not read the conversation

Problem
A messenger needs a server for delivery, and anything stored on a server can leak.
Choice
End-to-end encryption with MLS: the Rust library mls-rs, bridged to Android through uniffi.

The server only sees encrypted data. It orders group changes, deletes delivered messages and sends push notifications with no content. The local database on the phone is encrypted with SQLCipher.

Speech recognition without the cloud

Problem
Extracting the gist means recognizing speech. Sending audio to an external service would break end-to-end encryption.
Choice
Recognition and summarization run on the device.

Audio and text never leave the phone. I am currently improving the summarization model.