iTantra Download

Smart India Hackathon 2026 · ISRO · Problem Statement 26173

Speak in your language.
Send a few dozen bytes.
Be heard on the other side.

iTantra is an offline voice walkie-talkie for India. Speech becomes text on one phone, crosses Wi-Fi, Bluetooth LE or a LoRa radio bridge, and is spoken aloud on the other phone in 10 languages. No tower, no internet, no cloud.

Version
0.1.0 (debug build)
Needs
Android 8.0+, 64-bit ARM
Models
1.2 GB, loaded once (how)
SHA-256
ad15d2d37ef5933a70f0828266e03e155f68f2bb0ab3b8e559bfcdea2b06f5e0
iTantra talk screen: a red Hindi flood warning card, Talk, Call and Alert buttons, quick replies, SOS and the push-to-talk button
A flood alert from the control room, spoken aloud at full volume on arrival.
  • हिन्दी
  • বাংলা
  • मराठी
  • తెలుగు
  • தமிழ்
  • ગુજરાતી
  • ಕನ್ನಡ
  • മലയാളം
  • ଓଡ଼ିଆ
  • English

The problem

Voice is heavy. The links that survive a disaster are thin.

When towers fail after a flood or cyclone, what is left is Bluetooth, a crowded Wi-Fi hotspot, or a long-range radio that carries a few hundred bits per second. A spoken sentence needs 256 kbps as raw audio and still about 12 kbps with a good voice codec.

iTantra sends the meaning instead of the sound. It turns speech into text on the phone, sends the text, and speaks it again on arrival. The listener still hears a voice in their own language, so people who cannot read are included.

Bars use a log scale and assume a 3-second sentence. iTantra frame sizes are medians from a unit test over 310 FLEURS sentences, before encryption and link headers.

How it works

Five steps from mouth to ear

  1. SpeakHold the talk button. The phone listens for a pause of half a second to know a sentence has ended.
  2. RecogniseSraVaani, an open speech model from IISc, turns the sentence into text on the phone, about 20× faster than real time.
  3. PackThe text is compressed with a per-language dictionary into a small binary frame, encrypted and signed when phones are paired.
  4. SendThe frame goes over Wi-Fi, then Wi-Fi Direct, then Bluetooth LE, whichever works. A LoRa bridge extends it to kilometres.
  5. Speak againThe other phone voices the text with an Indian TTS model. Alerts jump the queue and play at full volume.

One message on the wire binary Msg frame, bytes

ver·type1 flags1 message id4 sender2 lang·emo1 TTL1 timestamp6 codec1 compressed text≈40–75

+ 11 B location on SOS+ 64 B Ed25519 signature on ALERT+ 20 B nonce and tag when encrypted

System diagrams

What the prototype does today

Features, with honest status

Each item was checked against the source code on 29 Sep 2026. Built works in the app now. Partial exists with gaps. Planned is not started.

Built

Push-to-talk

Hold to speak; the other phone sees you have the floor and holds its mic. Also on a Quick Settings tile.

Built

Call mode

Hands-free both ways with echo cancelling and noise suppression. Speakers take turns: the mic pauses while the phone talks.

Built

10 languages in and out

Hindi, Bengali, Marathi, Telugu, Tamil, Gujarati, Kannada, Malayalam, Odia and English, for recognition and speech. The app UI is translated too.

Built

Alerts that cannot be missed

ALERT messages use the alarm stream at full volume with exclusive audio focus and a vibration pattern.

Built

SOS with location

One tap sends one of 6 templates (rescue, medical, fire, flood, trapped, safe) with GPS/NavIC coordinates, read aloud on arrival.

Built

Live captions

Partial text streams to the listener every half second while the speaker is still talking.

Built

Three-link fallback

Wi-Fi LAN first; Wi-Fi Direct after 4 s and Bluetooth LE after 8 s if there is no Wi-Fi. Same frames on every link.

Built

Compact binary frames

Negotiated between phones on connect. Median 62 B per sentence in a unit test, against 524 B for JSON.

Built

LITE profile for smaller phones

Chosen automatically under 6 GB RAM: 2 threads, no live captions, voice engines unload after 60 s idle.

Built

Gapless playback

A producer/consumer audio queue with adaptive pre-roll. 0 underruns across 5 languages on a Snapdragon 870.

Built

Metrics, history, replay

Latency p50/p95, real-time factor, CPU, memory and thermal state on screen; searchable history; replay the last 20 voice notes.

Built

Control-room dashboard

A Python gateway on a laptop joins the network, logs every message, shows peers and latency, broadcasts alerts and reads CAP 1.2 files.

Partial

Encryption and signed alerts

X25519 pairing by QR code, ChaCha20-Poly1305, Ed25519-signed ALERTs. Covers binary mode after pairing; captions and the JSON path are not covered yet.

Partial

ESP32 LoRa bridge

Firmware for ESP32-S3 + SX1262 at 865 MHz compiles (RAM 9.3%, flash 17.2%). Not yet flashed or range-tested.

Partial

Mesh relay

The bridge firmware relays with a hop limit and duplicate filter. Phones do not relay yet.

Planned

Translation on device

A Hindi speaker heard in Tamil, using IndicTrans2. Also planned: small VITS voices for hi/gu/or/en, sender-voice rendering, wake word.

See it working

Demo video and phone recordings

Narrated walkthrough of the whole system, 1 min 58 s, with English captions.
Receive and speak Messages in Hindi, Tamil and English arrive and are spoken aloud. 25 s
Flood alert A Hindi ALERT arrives, flashes red and plays at alarm volume. 16 s
SOS One tap sends "बचाव चाहिए" (need rescue) with GPS coordinates. 18 s
Speech to text Hindi and Tamil test recordings through the real on-device recogniser (audio injected, not live mic). 20 s
First run Greeting in every script, language picker, offline promise, credits. 30 s
Metrics and history CPU, RAM, latency histogram, real-time factor, searchable log. 24 s
Hindi interface Switching the whole app to हिन्दी and back. 38 s
Settings Text size S to XL, dark and light theme, speech speed 0.8× to 1.3×. 32 s

Measured on real phones

Prototype numbers

From the app's own logging on a Snapdragon 870 (vivo I2202) and a Snapdragon 8 Gen 3 (OnePlus 12), unless the note says otherwise.

0.05speech-recognition real-time factor, about 20× faster than speech
41 msmessage to acknowledgement over Wi-Fi
0.3–1.6%idle CPU in push-to-talk mode
688 MBmemory in the LITE profile when idle
MetricValueNote
Speech-recognition word error rate5.9%10 FLEURS clips, 1 per language. Preliminary; a 30-clip-per-language run is next
Speech end → text ready0.4–0.9 sSnapdragon 870
Text-to-speech first audio0.6–0.95 sIndic-Mio (hi/gu/en); VITS languages start faster
Audio underruns during playback0hi, en, gu, ta, bn on the Snapdragon 870
Alert on the alarm stream16/16Full volume with exclusive audio focus
End to end, said → heard≈1.5–2.5 sEstimate from one phone in both roles with a peer simulator
Median bytes per sentence62 BJVM unit test, 310 sentences; the binary format itself runs in the app
APK size60 MBNative runtimes; models are installed separately

A hot phone can run 2 to 4 times slower. The Metrics screen shows thermal state beside every number, so results can be compared fairly.

Open source, end to end

Models and technology

SraVaani-1.0

ARTPARK, IISc Bengaluru · MIT

Speech recognition for all 10 languages. FastConformer, int8 ONNX, 487 MB.

Indic-Mio

SPRING Lab, IIT Madras · Apache-2.0

Expressive speech for hi, gu, or, en. 0.6B, Q4 GGUF via llama.cpp, 392 MB + 194 MB codec.

VITS Rasa

AI4Bharat, IIT Madras · CC-BY-4.0

Fast speech for bn, kn, ml, mr, ta, te with 14 emotion styles. 123 MB.

Silero VAD

Silero · MIT

Detects when someone is speaking and when a sentence ends. 0.6 MB.

LayerTechnology
AppKotlin 2.4, Jetpack Compose, Material 3, foreground service; Android 8.0+ (API 26), arm64
Speech runtimesherpa-onnx 1.13.8 (ONNX Runtime), llama.cpp / ggml with ARM dot-product kernels, C++17 over JNI
TextIndic number, date, currency and coordinate normaliser; SSML subset; alert-keyword style detection
TransportTCP over Wi-Fi with mDNS discovery, Wi-Fi Direct, BLE GATT (Nordic UART); binary frames with per-script dictionaries
SecurityBouncyCastle: X25519, HKDF-SHA256, ChaCha20-Poly1305, Ed25519; Android Keystore; QR pairing with ZXing and CameraX
HardwareESP32-S3 + SX1262 LoRa (865–867 MHz licence-free band), NimBLE, RadioLib; about ₹1,900–3,200 per node
Control roomPython gateway, SQLite log, live dashboard with server-sent events, CAP 1.2 alert parser

For reviewers

Install and test the prototype

About 15 minutes the first time. You need an Android phone (8.0 or newer, 64-bit, 6 GB RAM or more recommended, about 2 GB free) and, for the speech models, a computer with a USB cable.

  1. Install the app

    On the phone, open this page and tap Download APK. Open the file and allow Install unknown apps for your browser when Android asks. This is a debug build signed with a development key, so Play Protect may warn you; choose Install anyway.

    Open iTantra once and grant the microphone and notification permissions.

  2. Load the speech models (1.2 GB, once)

    The models are too large to put inside the APK, so they are copied to the phone over USB. Turn on Developer options → USB debugging, connect the phone, and accept the prompt. You also need Android platform-tools (adb).

    curl -fsSL https://itantra-106.pages.dev/install-models.sh -o install-models.sh
    bash install-models.sh

    The scripts check every file's checksum and skip files already downloaded, so you can safely run them again. Read them first: install-models.sh · install-models.ps1.

  3. Check the models

    Open iTantra → Settings → Models. Every row should show as present. Without the models the app still opens and you can explore every screen, but it cannot recognise or speak.

  4. Talk between two phones

    Install on two phones and put both on the same Wi-Fi (a phone hotspot works). They find each other automatically within a few seconds; the link type shows at the top. With no Wi-Fi they try Wi-Fi Direct, then Bluetooth.

    Only one phone? Run the control-room gateway from the source code (tools/gateway) on a laptop on the same Wi-Fi and send messages from its dashboard.

  5. Things to try

    • Pick your language, hold the orange button, say a sentence, release. Watch the text appear and the other phone speak it.
    • Turn on Alert and send "flood water is rising". The other phone plays it on the alarm channel at full volume and vibrates.
    • Tap SOS and choose a template. The location is spoken on the other phone.
    • Switch to Call for hands-free conversation.
    • Open Metrics to see latency, real-time factor, memory and thermal state.

Known limitations

  • Debug build (0.1.0): expect rough edges. Tested on Snapdragon 870 and 8 Gen 3 phones.
  • The first message after opening can take a few seconds while a voice model loads.
  • Recognition accuracy drops in noisy places and with mixed languages.
  • The LoRa bridge and phone-to-phone relaying are not part of this build's test.

What comes next

Roadmap

  1. Working prototype10-language speech in and out, alerts, SOS, three links, binary frames, encryption, metrics, gateway, demo video.
  2. Idea round30-clip accuracy per language, speech intelligibility test, end-to-end latency p50/p95, a 4 GB phone test.
  3. Before the finaleSmall fast voices for hi/gu/or/en, smaller recogniser for low-end phones, LoRa bridge on real hardware, two-phone field test.
  4. FinaleLive multi-hop LoRa demo in two languages, phone relaying, place names for SOS, listener study.
  5. AfterOn-device translation, the sender's own voice, live NDMA SACHET alerts, pilot with a district disaster team.

Credits and references

Built on Indian open AI research