Live English transcription and multilingual subtitles for Devcon 8, on infrastructure paid in ETH

Live English transcription and multilingual subtitles for Devcon 8, on infrastructure paid in ETH

Devcon 8 is in Mumbai, and I’d like to offer something for it: a live English transcript of every talk, readable on your own phone, with translation into Indic and other languages offered on top for anyone who wants it. It’s generated by open-source models running on community-operated infrastructure — where the people running that infrastructure are paid in ETH on Arbitrum, per second of audio, with no contract and no invoice.

I’m putting this here for discussion before submitting it as a DIP. I’d rather find out now if it’s daft.

Updated since first posting: this originally led with translation and treated the English transcript as the plumbing underneath it. That was backwards, and Dr. Jeevan Deol, reviewing real output as a native Hindi speaker, is the reason I know. See the comment below for what changed and why.

Draft DIP metadata
Title: Live English transcription and multilingual subtitles, served by community infrastructure paid in ETH
Status: Draft
Themes: Accessibility, Local Impact, Infrastructure
Tags: accessibility, av, transcription, translation, india, open-source, depin
Instances: Devcon 8 (Mumbai)
Resources Required: an AV audio feed per stage

What it does

Speech from a talk is transcribed into English as it is spoken, and optionally translated into several languages at once. Each subtitle carries the timestamps of the audio it describes, along with every translation of that same interval, so each person reads what they want on their own phone, in step with the speaker.

Eight Indian languages are working today — Hindi, Marathi, Gujarati, Punjabi, Tamil, Telugu, Malayalam and Kannada — via IndicTrans2, which covers 22 Indic languages from the same model, so the rest are configuration rather than new work. Each is available in its own script and in Latin letters, which turns out to matter more than I expected. Seventeen further language pairs are benchmarked and working.

If the event also wants transcoded video renditions, it produces those from the same pass over the stream, since the audio only has to be decoded once. The code is public domain.

Why the English transcript is the main event

This is the part I got wrong first time round.

Most attendees at an event like this are entirely comfortable in English. A smaller group learnt English at school, got into tech, built up their English along the way — and still lose the occasional sentence to a fast speaker or an unfamiliar accent. That group, in Dr. Deol’s words, would be “very relieved by an English transcription” and might occasionally glance at a translation to get the gist.

That’s a much better person to design for than an abstract non-English speaker. And the same transcript serves deaf and hard-of-hearing attendees, anyone at the back of a hall, anyone whose attention wandered, and everyone watching the recording afterwards.

It’s also the half that works best, costs least to run, and has the lowest latency. Translation is worth offering and worth doing well. It just isn’t the product.

The language question in India isn’t a technical question

I’d have got this wrong without asking, and I think it’s worth flagging to Devcon regardless of whether any of my software gets used, because it applies to any translated content shown at the event.

There is no single language that can go on a shared screen in India without being a political statement.

  • Hindi is a northern consensus, not an Indian one.
  • Mumbai is a Marathi-speaking city, and Marathi is Maharashtra’s state language. Including it is ordinary courtesy towards the place hosting the event, and its absence - particularly alongside a Hindi track - would be noticed. Language is a charged subject in Maharashtra, which seems to me a reason to be thoughtful about it rather than a reason to take anybody’s side.
  • South India broadly dislikes Hindi, and there’s no southern lingua franca and no shared southern script. Tamil, Telugu, Malayalam and Kannada are four languages with four writing systems and no common fallback. Southern attendees would generally prefer English.
  • People differ on alphabet, not just language. Indians overwhelmingly text in Latin letters, and many read romanised Hindi faster than Devanagari — while attendees from smaller cities and Hindi-speaking regions may well prefer Devanagari.

This is the strongest argument for the approach here, and I’d say it even if someone else were building it. Because everyone reads on their own phone and picks their own language and their own script, nobody has to agree and nothing is imposed on a room. A shared caption screen would force Devcon into a language-politics decision it has no reason to make. This removes that decision entirely — and it’s also why carrying eight or more Indian languages is proportionate rather than showing off. There is no smaller set that works.

Where it would actually run

On Livepeer orchestrators with GPUs — of which the network has a great deal, already serving other AI workloads.

I built and proved this on deliberately modest hardware, and I mention that only for what it says about the floor: the whole pipeline — transcription plus six simultaneous translations — fits on a single-board computer drawing eight watts. That is a statement about how little this costs to run, not a proposal for how to run a conference. Production belongs on GPU orchestrators, where the headroom buys more languages, lower latency, and redundancy across independent operators.

This is where I need the Livepeer community. The capability is written and published; what it needs is operators willing to run it for the week, ideally two or three independently, so that no single machine is the reason subtitles stop mid-talk. I’ve started that conversation over on the Livepeer forum — Live English→Hindi captions for Devcon 8, served by Livepeer orchestrators: floating the idea - Livepeer Forum — and I’d rather arrive there with Devcon’s interest than the other way round.

The part I’d most like you to look at

This is community-run, sovereign media infrastructure, and the people running it are paid in ETH on Arbitrum, an Ethereum L2 — metered per second of audio, settled in probabilistic micropayments, with no contract, no invoice, no purchase order, and no account with anybody. An operator advertises a capability, a client discovers them and pays them, and that is the entire commercial relationship.

It is Ethereum used as ordinary infrastructure plumbing, at an event about Ethereum, doing something the event actually needs. Devcon would be the first event to procure infrastructure this way.

Why not just use a cloud API

Being straightforward, because someone will ask.

For English transcription — the primary deliverable — open models are genuinely strong, and I don’t think there’s much of a gap to concede. For translation, I’d assumed a large commercial provider would beat this comfortably, particularly on Indian languages. I’m now less sure than I was: Dr. Deol, having used both, told me that mainstream commercial voice translation “struggles a lot with Indian languages”, and that our output was correct, if stilted. That’s one person’s read on one talk and I’m not going to over-claim from it, but I’m no longer willing to state the opposite as fact either.

The other arguments are unchanged:

  • It is free and open source, public domain, and it does not expire. Devcon can keep it, change it, or hand it to the next event.
  • No vendor lock-in and no data processor. Audio from the stage is handled by infrastructure the community runs, under a licence anyone can read, rather than sent to a commercial API under terms nobody does.
  • The operators are independent, so if one fails another serves the session.
  • The cost is small, metered, and paid to whoever provides the compute.

And it can run alongside a commercial service rather than instead of it.

What is not yet proven

I’d rather name these now than in November.

Indic translation quality is under review right now, by people who can actually read it. This has moved on since I first posted: IndicTrans2 (AI4Bharat, MIT-licensed) is integrated and running, and eight Devcon SEA talks translated into eight Indian languages are with native speakers. The early signal is that the translations are correct — better than expected — but that they read as “translaterese”: formally correct Hindi where a real speaker would drop English words into a Hindi sentence. The example Dr. Deol gave of how someone actually talks: “mai b sochta hu k apple users uska interface prefer karenge”. I won’t put Indic subtitles on a screen in Mumbai on the strength of a metric.

Technical and everyday vocabulary is the specific weakness, and it’s tractable. Measured, not assumed: ordinary speech translates well and the vocabulary of an Ethereum talk does not. The clearest example is orchestrator, which the European models return as a musical orchestra in Dutch, Swedish, Russian, Chinese and Indonesian; “stream” becomes electricity in German and Spanish, and a brook in Czech. Only Italian and Romanian got it right.

The fix is a glossary of terms that should be left in English or translated one agreed way, in three categories: Ethereum terms, general technical terms, and ordinary functional words that Indian-language speakers simply say in English. One-off work that improves every language at once, and a good task for people who already translate Ethereum content. A draft is circulating.

IndicTrans2 often fails more gently than the European models: much of the time it transliterates technical English into the native script rather than mistranslating it, and where that happens the fix is just restoring the English spelling in the romanised output - a lookup table rather than a model change. Measured on all eight languages, orchestrator survives as a transliteration in six and becomes a musical ensemble in Tamil and Kannada.

But it is not the general case, and the counter-example matters more than the example. “Wallet” comes back as a physical purse or handbag in seven of the eight, and only Telugu keeps the English. That is a core Ethereum word failing quietly into a plausible wrong meaning, which is worse than failing loudly. So the glossary has to do the harder job - forcing the English through the translation - and not merely tidy up the spelling afterwards.

Nobody has run this at event scale, with event audio, in a hall.

Things that would make this much better

The Bhasha CROPS Community Hub. Their proposal includes live human translation of a main-stage talk into Indian languages. That is the same goal from the opposite direction — they have the languages, the native speakers and the standing in those communities. Live human interpretation is better than any machine, but it does not scale to every talk on every day, and it does not carry over to the recordings. This does. I have posted there separately and would like to propose these together.

Crucially: they should be the ones who decide whether the output is good enough, because I cannot read it. Machine translation without native-speaker review is how you end up with confident nonsense on a screen in front of the people it was meant to serve. Everything in the language-politics section above came from asking Dr. Jeevan Deol over the course of a single afternoon, which is an argument for asking a great many more people, early.

Whatever happens with streaming. If some other pipeline ends up decoding every stage’s audio anyway, transcription and translation come nearly free off the back of that — captions are not something a streaming stack usually arrives with. Worth a conversation with whoever ends up doing it.

The ethereum.org translation programme. Roughly 2.89 million words across 68 languages, translated by humans who understood the subject. That is precisely the parallel text that would fix the vocabulary problem, either as a fine-tuning corpus or, far more cheaply, as the source of the glossary. The programme is winding down, which to me is an argument for making sure the corpus outlives it. Open question I can’t answer from outside: is that translation memory exportable and openly licensed? If someone knows, it changes what is possible here.

What I’d need from Devcon

To get started:

  • An audio feed from one or more talk rooms — SRT, RTMP, or a line-level tap into a machine I provide. Plus a conversation with whoever owns the AV workstream, early enough to be useful.

I’ll build the rest, including where it gets displayed. The plan is a page per stage on an ENS name — open it on your phone, pick your language and your script, read along. It scales to any number of languages because each person picks their own, it needs no screen real estate and no argument about what goes on a confidence monitor, and as set out above it imposes no language decision on a room. Hosted on decentralised storage, so the compute, the hosting and the naming are all community infrastructure rather than someone’s account.

Deliberately text only, which matters more on site than it sounds: subtitles are a few kilobytes a minute, video is megabits, and a hall full of people all pulling a stream to read along would be a poor way to spend the venue’s wifi. Anyone watching the livestream can of course have the subtitles rendered there instead — that is a decision for whoever runs the stream, and the captions are just timestamped text, so it is not a difficult one.

One thing I’d flag honestly, because it changes what “good” means: in the room you cannot delay reality. For a stream or a recording, playback can be held back so subtitles land exactly on the words they describe. Someone sitting in the hall hears the speaker live, so any latency is felt directly as text trailing the voice. Human simultaneous interpreters typically trail by two to six seconds, and that is the bar worth hitting. Transcription alone is comfortably faster than transcription plus translation — a further practical reason for the English transcript being the primary track, with translations trailing slightly further behind for those who opt in.

The useful replies

  • “I’d run an orchestrator for it” — especially with a GPU
  • “I speak one of these languages and I’ll review output” — the single most useful thing anyone can offer
  • “Here’s why the model choice is wrong”
  • “This is daft, because…” — genuinely, better now than in Mumbai.
2 Likes

Update: I put this in front of Dr. Jeevan Deol, who reviewed it as a native Hindi speaker, and it changed the proposal. I’ve edited the post above; here’s what changed and why.

The short version is that I had it backwards. I was offering multilingual translation, with the English transcript as the plumbing that made translation possible. The first point back was that most Indians at an event like Devcon aren’t looking for a translation at all — they want a transcription into English, in case a speaker’s accent or speed is hard to grasp. Translation is worth offering, but as an option rather than the product.

Dr. Deol also described the person it helps far better than I had: someone who learnt English at school, got into tech, built their English up along the way, and is entirely competent but still loses the occasional sentence. That person would be “very relieved by an English transcription”. I’ve rewritten the proposal around them.

The part I most want to pass on to Devcon has nothing to do with my software.

There is no single language that can go on a shared screen in India without being a political statement. Hindi is a northern consensus rather than an Indian one. In Mumbai specifically it’s loaded — Maharashtra’s politics run through Marathi identity, and Hindi on a screen without Marathi can read as a slight. And South India broadly dislikes Hindi, with no southern lingua franca and no shared southern script; Tamil, Telugu, Malayalam and Kannada are four languages with four writing systems and no common fallback.

I’d have walked straight into that. Whoever ends up handling translated content at Devcon 8 should know it whether or not any of this proposal goes anywhere.

It also settled a design question I’d been going back and forth on. Everyone reads on their own phone, picks their own language, and — this was the bit I hadn’t anticipated — picks their own alphabet. Indians overwhelmingly text in Latin letters and many read romanised Hindi faster than Devanagari, while attendees from smaller cities and Hindi-speaking regions often prefer Devanagari. So both, and let people choose. That costs a toggle rather than a second translation, because the Latin version is a mechanical conversion of the native-script output.

On quality, two corrections to what I wrote before:

I’d said the freely available Hindi models weren’t good enough for a stage and that IndicTrans2 was a promising direction. IndicTrans2 is now built, running, and eight Devcon SEA talks are translated into eight Indian languages and sitting with native speakers. The early read is that the translations are correct — the problem is that they’re stilted. Formally correct Hindi where a real person would drop English words into a Hindi sentence: “mai b sochta hu k apple users uska interface prefer karenge” is what natural speech looks like.

And I’d conceded that a commercial API would probably beat this on Hindi. Dr. Deol, having used both, told me that mainstream commercial voice translation “struggles a lot with Indian languages” and that this output was better than expected. That’s one person on one talk and I’m not going to over-claim from it — but I’m no longer prepared to assert the opposite either, so I’ve softened it.

The glossary got sharper too. It needs three categories rather than one: Ethereum terms, general technical terms, and ordinary functional words that Indian-language speakers just say in English. And there’s a pleasant surprise in the detail — IndicTrans2 mostly transliterates technical English into Devanagari rather than mistranslating it, so a good part of the fix is restoring the English spelling in the romanised output. A lookup table rather than a model change. orkestretar vidiyo ko netavark par ek nod par strim karta hai becomes orchestrator video ko network par ek node par stream karta hai.

What I’d ask for. All of the above came from one person over one afternoon, and Dr. Deol hasn’t finished reading the sample talk yet. If you speak any Indian language and would look at a talk’s worth of output, that is worth more to this proposal than anything else anyone can offer — including a GPU.

1 Like