Push-to-talk
Hold to speak; the other phone sees you have the floor and holds its mic. Also on a Quick Settings tile.
Smart India Hackathon 2026 · ISRO · Problem Statement 26173
iTantra is an offline voice walkie-talkie for India. Speech becomes text on one phone, crosses Wi-Fi, Bluetooth LE or a LoRa radio bridge, and is spoken aloud on the other phone in 10 languages. No tower, no internet, no cloud.
ad15d2d37ef5933a70f0828266e03e155f68f2bb0ab3b8e559bfcdea2b06f5e0
The problem
When towers fail after a flood or cyclone, what is left is Bluetooth, a crowded Wi-Fi hotspot, or a long-range radio that carries a few hundred bits per second. A spoken sentence needs 256 kbps as raw audio and still about 12 kbps with a good voice codec.
iTantra sends the meaning instead of the sound. It turns speech into text on the phone, sends the text, and speaks it again on arrival. The listener still hears a voice in their own language, so people who cannot read are included.
Bars use a log scale and assume a 3-second sentence. iTantra frame sizes are medians from a unit test over 310 FLEURS sentences, before encryption and link headers.
How it works
One message on the wire binary Msg frame, bytes
+ 11 B location on SOS+ 64 B Ed25519 signature on ALERT+ 20 B nonce and tag when encrypted
What the prototype does today
Each item was checked against the source code on 29 Sep 2026. Built works in the app now. Partial exists with gaps. Planned is not started.
Hold to speak; the other phone sees you have the floor and holds its mic. Also on a Quick Settings tile.
Hands-free both ways with echo cancelling and noise suppression. Speakers take turns: the mic pauses while the phone talks.
Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam, Odia and English, for recognition and speech. The app UI is translated too.
ALERT messages use the alarm stream at full volume with exclusive audio focus and a vibration pattern.
One tap sends one of 6 templates (rescue, medical, fire, flood, trapped, safe) with GPS/NavIC coordinates, read aloud on arrival.
Partial text streams to the listener every half second while the speaker is still talking.
Wi-Fi LAN first; Wi-Fi Direct after 4 s and Bluetooth LE after 8 s if there is no Wi-Fi. Same frames on every link.
Negotiated between phones on connect. Median 62 B per sentence in a unit test, against 524 B for JSON.
Chosen automatically under 6 GB RAM: 2 threads, no live captions, voice engines unload after 60 s idle.
A producer/consumer audio queue with adaptive pre-roll. 0 underruns across 5 languages on a Snapdragon 870.
Latency p50/p95, real-time factor, CPU, memory and thermal state on screen; searchable history; replay the last 20 voice notes.
A Python gateway on a laptop joins the network, logs every message, shows peers and latency, broadcasts alerts and reads CAP 1.2 files.
X25519 pairing by QR code, ChaCha20-Poly1305, Ed25519-signed ALERTs. Covers binary mode after pairing; captions and the JSON path are not covered yet.
Firmware for ESP32-S3 + SX1262 at 865 MHz compiles (RAM 9.3%, flash 17.2%). Not yet flashed or range-tested.
The bridge firmware relays with a hop limit and duplicate filter. Phones do not relay yet.
A Hindi speaker heard in Tamil, using IndicTrans2. Also planned: small VITS voices for hi/gu/or/en, sender-voice rendering, wake word.
Real screenshots
Captured from the running prototype. Tap any screen to see it full size.
See it working
Measured on real phones
From the app's own logging on a Snapdragon 870 (vivo I2202) and a Snapdragon 8 Gen 3 (OnePlus 12), unless the note says otherwise.
| Metric | Value | Note |
|---|---|---|
| Speech-recognition word error rate | 5.9% | 10 FLEURS clips, 1 per language. Preliminary; a 30-clip-per-language run is next |
| Speech end → text ready | 0.4–0.9 s | Snapdragon 870 |
| Text-to-speech first audio | 0.6–0.95 s | Indic-Mio (hi/gu/en); VITS languages start faster |
| Audio underruns during playback | 0 | hi, en, gu, ta, bn on the Snapdragon 870 |
| Alert on the alarm stream | 16/16 | Full volume with exclusive audio focus |
| End to end, said → heard | ≈1.5–2.5 s | Estimate from one phone in both roles with a peer simulator |
| Median bytes per sentence | 62 B | JVM unit test, 310 sentences; the binary format itself runs in the app |
| APK size | 60 MB | Native runtimes; models are installed separately |
A hot phone can run 2 to 4 times slower. The Metrics screen shows thermal state beside every number, so results can be compared fairly.
Open source, end to end
ARTPARK, IISc Bengaluru · MIT
Speech recognition for all 10 languages. FastConformer, int8 ONNX, 487 MB.
SPRING Lab, IIT Madras · Apache-2.0
Expressive speech for hi, gu, or, en. 0.6B, Q4 GGUF via llama.cpp, 392 MB + 194 MB codec.
AI4Bharat, IIT Madras · CC-BY-4.0
Fast speech for bn, kn, ml, mr, ta, te with 14 emotion styles. 123 MB.
Silero · MIT
Detects when someone is speaking and when a sentence ends. 0.6 MB.
| Layer | Technology |
|---|---|
| App | Kotlin 2.4, Jetpack Compose, Material 3, foreground service; Android 8.0+ (API 26), arm64 |
| Speech runtime | sherpa-onnx 1.13.8 (ONNX Runtime), llama.cpp / ggml with ARM dot-product kernels, C++17 over JNI |
| Text | Indic number, date, currency and coordinate normaliser; SSML subset; alert-keyword style detection |
| Transport | TCP over Wi-Fi with mDNS discovery, Wi-Fi Direct, BLE GATT (Nordic UART); binary frames with per-script dictionaries |
| Security | BouncyCastle: X25519, HKDF-SHA256, ChaCha20-Poly1305, Ed25519; Android Keystore; QR pairing with ZXing and CameraX |
| Hardware | ESP32-S3 + SX1262 LoRa (865–867 MHz licence-free band), NimBLE, RadioLib; about ₹1,900–3,200 per node |
| Control room | Python gateway, SQLite log, live dashboard with server-sent events, CAP 1.2 alert parser |
For reviewers
About 15 minutes the first time. You need an Android phone (8.0 or newer, 64-bit, 6 GB RAM or more recommended, about 2 GB free) and, for the speech models, a computer with a USB cable.
On the phone, open this page and tap Download APK. Open the file and allow Install unknown apps for your browser when Android asks. This is a debug build signed with a development key, so Play Protect may warn you; choose Install anyway.
Open iTantra once and grant the microphone and notification permissions.
The models are too large to put inside the APK, so they are copied to the phone over USB. Turn on Developer options → USB debugging, connect the phone, and accept the prompt. You also need Android platform-tools (adb).
curl -fsSL https://itantra-106.pages.dev/install-models.sh -o install-models.sh
bash install-models.sh
curl.exe -fsSL https://itantra-106.pages.dev/install-models.ps1 -o install-models.ps1
powershell -ExecutionPolicy Bypass -File install-models.ps1
The script reads the model list, downloads each file from …/models/v1/<path>, joins the files that are split into .part-aa, .part-ab… pieces, checks the SHA-256 and runs:
adb shell mkdir -p /sdcard/Android/data/org.itantra.app/files/models
adb push vad stt tts /sdcard/Android/data/org.itantra.app/files/models/
The scripts check every file's checksum and skip files already downloaded, so you can safely run them again. Read them first: install-models.sh · install-models.ps1.
Open iTantra → Settings → Models. Every row should show as present. Without the models the app still opens and you can explore every screen, but it cannot recognise or speak.
Install on two phones and put both on the same Wi-Fi (a phone hotspot works). They find each other automatically within a few seconds; the link type shows at the top. With no Wi-Fi they try Wi-Fi Direct, then Bluetooth.
Only one phone? Run the control-room gateway from the source code (tools/gateway) on a laptop on the same Wi-Fi and send messages from its dashboard.
What comes next
Credits and references