Offline Real-Time Multilingual Broadcasting Using WebRTC and Local AI Services: Architecture, Performance, and Institutional Evaluation
Abstract
1. Introduction
- A fully offline multilingual broadcasting architecture operating within institution-owned infrastructure.
- A custom-fork WebRTC SFU extended with language-specific media routing, per-language RTP audio management, and recording logic.
- A local AI pipeline integrating GGML large-v3-turbo speech-to-text, a quantized neural translation model, and local TTS voice models.
- A multilingual delivery model in which original and translated audio are carried as separate, selectable channels.
- A browser/PWA viewer with QR-code onboarding, WebRTC primary delivery, and HLS fallback.
- A security and access-control design combining token-based access, JWT admin login, license validation, IP-to-room assignment, and university room/card-service integration.
- A practical institutional evaluation covering latency budget, bandwidth estimation, pilot stakeholder feedback, data sovereignty, and total-cost-of-ownership implications.
- RQ1. Can a fully offline WebRTC-SFU-based multilingual broadcasting platform support real-time classroom and presentation scenarios without cloud services?
- RQ2. What is the end-to-end latency budget of the local STT → translation → TTS → WebRTC delivery pipeline, and which components dominate perceived latency?
- RQ3. How does the system support multilingual delivery through per-language audio tracks while minimizing unnecessary media delivery to viewers?
- RQ4. What are the bandwidth implications of delivering video, original audio, and multiple synthesized translation channels in representative classroom scenarios?
- RQ5. What authentication, authorization, recording, transcript, and institutional integration mechanisms are required for real-world university deployment?
- RQ6. What are the institutional implications of offline deployment in terms of privacy, data sovereignty, accessibility, cost, and stakeholder acceptance?
2. Related Work and Background
2.1. Speech Recognition and Local Inference
2.2. Neural Machine Translation and Quantized Local Models
2.3. Neural Speech Synthesis
2.4. Institutional Implications of Offline AI
3. System Architecture
3.1. Ingest and Transport
3.2. Publish Modes
3.3. AI Processing Tier
3.4. Distribution and Client Access
3.5. Recording, Transcript, and Archive
3.6. Security and Access Control
4. The AI Pipeline
4.1. Voice Activity Detection and Speech-to-Text
4.2. Translation Layer
4.3. Translation Context and Terminology
4.4. Text-to-Speech Synthesis
4.5. Per-Language Audio Track Management
4.6. Service Orchestration and Internal Messaging
5. Methodology
5.1. Deployment Environment
5.2. Pilot Evaluation and Data Collection
6. Technical Results
6.1. Latency Budget Analysis
6.2. End-to-End Latency Estimation
6.3. Language Coverage
6.4. Bandwidth Estimation
6.5. Component Model Accuracy (Published Benchmarks)
6.6. Projected Scalability and GPU Utilization
6.7. Pilot Deployment Observations
7. Institutional Impact Analysis
7.1. Total Cost of Ownership
7.2. Accessibility and Universal Design for Learning
7.3. Data Sovereignty, Regulatory Compliance, and Connectivity
7.4. Pilot Demonstrations and Stakeholder Feedback
8. Discussion
8.1. Threats to Validity
8.2. Future Work
- Controlled, deployment-specific WER evaluation on Turkish academic speech (beyond the published model-level WER of ≈7.8–8.4%).
- MetricX/COMET, ChrF, and BLEU evaluation on institutional content for the eight active target languages.
- Human translation-quality evaluation with qualified raters.
- Measured GPU-utilization profiling on the RTX 5090 deployment target under multi-broadcast load (beyond the estimated 20–25% single-broadcast footprint).
- Benchmarking across lower-end GPU tiers to quantify the trade-offs among VRAM capacity, inference latency, throughput, and deployment cost.
- Multi-session concurrency benchmarking under controlled load.
- Empirical scalability testing at 10, 50, and 100 simultaneous listeners to validate the bandwidth projections of Table 6.
- Network stress testing under varied bandwidth and loss conditions.
- Robustness evaluation under speech noise, packet loss, bandwidth degradation, and long-utterance conditions.
- Fault-injection experiments involving temporary STT, translation, TTS, and WebRTC service interruptions to quantify recovery behavior and service continuity.
9. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| GDPR | General Data Protection Regulation |
| GGML | GGML model format |
| HLS | HTTP Live Streaming |
| JWT | JSON Web Token |
| KVKK | Personal Data Protection Law (Türkiye) |
| PWA | Progressive Web Application |
| RMS | Root Mean Square |
| RTP | Real-time Transport Protocol |
| SFU | Selective Forwarding Unit |
| STT | Speech-to-Text |
| TCO | Total Cost of Ownership |
| TTS | Text-to-Speech |
| UDL | Universal Design for Learning |
| VAD | Voice Activity Detection |
| WebRTC | Web Real-Time Communication |
References
- Radford, A.; Kim, J.W.; Xu, T.; Brockman, G.; McLeavey, C.; Sutskever, I. Robust speech recognition via large-scale weak supervision. In Proceedings of the International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems 30 (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar] [CrossRef] [Scilit]
- van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; Kavukcuoglu, K. WaveNet: A Generative Model for Raw Audio. In Proceedings of the 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), Sunnyvale, CA, USA, 13–15 September 2016; p. 125. [Google Scholar] [CrossRef] [Scilit]
- Altbach, P.G.; Knight, J. The internationalization of higher education: Motivations and realities. J. Stud. Int. Educ. 2007, 11, 290–305. [Google Scholar] [CrossRef] [Scilit]
- Knight, J. Internationalization remodeled: Definition, approaches, and rationales. J. Stud. Int. Educ. 2004, 8, 5–31. [Google Scholar] [CrossRef] [Scilit]
- Rabiner, L.R. A tutorial on hidden Markov models and selected applications in speech recognition. Proc. IEEE 1989, 77, 257–286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hinton, G.; Deng, L.; Yu, D.; Dahl, G.; Mohamed, A.; Jaitly, N.; Senior, A.; Vanhoucke, V.; Nguyen, P.; Sainath, T.; et al. Deep neural networks for acoustic modeling in speech recognition. IEEE Signal Process. Mag. 2012, 29, 82–97. [Google Scholar] [CrossRef] [Scilit]
- Graves, A.; Fernández, S.; Gomez, F.; Schmidhuber, J. Connectionist temporal classification. In Proceedings of the 23rd International Conference on Machine Learning (ICML), Pittsburgh, PA, USA, 25–29 June 2006; pp. 369–376. [Google Scholar] [CrossRef] [Scilit]
- Chan, W.; Jaitly, N.; Le, Q.; Vinyals, O. Listen, attend and spell: A neural network for large vocabulary conversational speech recognition. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China, 20–25 March 2016; pp. 4960–4964. [Google Scholar] [CrossRef] [Scilit]
- Junczys-Dowmunt, M.; Grundkiewicz, R.; Dwojak, T.; Hoang, H.; Heafield, K.; Neckermann, T.; Seide, F.; Germann, U.; Fikri Aji, A.; Bogoychev, N.; et al. Marian: Fast neural machine translation in C++. In Proceedings of the ACL 2018 System Demonstrations, Melbourne, Australia, 15–20 July 2018; pp. 116–121. [Google Scholar] [CrossRef] [Scilit]
- Scaling neural machine translation to 200 languages. Nature 2024, 630, 841–846. [CrossRef] [Scilit] [PubMed]
- Shen, J.; Pang, R.; Weiss, R.J.; Schuster, M.; Jaitly, N.; Yang, Z.; Chen, Z.; Zhang, Y.; Wang, Y.; Skerrv-Ryan, R.; et al. Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada, 15–20 April 2018; pp. 4779–4783. [Google Scholar] [CrossRef] [Scilit]
- Ren, Y.; Ruan, Y.; Tan, X.; Qin, T.; Zhao, S.; Zhao, Z.; Liu, T.-Y. FastSpeech: Fast, Robust and Controllable Text to Speech. In Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS 2019), Vancouver, BC, Canada, 8–14 December 2019; pp. 3165–3174. Available online: https://papers.neurips.cc/paper/2019/hash/f63f65b503e22cb970527f23c9ad7db1-Abstract.html (accessed on 15 June 2026).
- Couture, S.; Toupin, S. What does the notion of “sovereignty” mean when referring to the digital? New Media Soc. 2019, 21, 2305–2322. [Google Scholar] [CrossRef] [Scilit]
- Floridi, L. The fight for digital sovereignty: What it is, and why it matters, especially for the EU. Philos. Technol. 2020, 33, 369–378. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Selwyn, N. Education and Technology: Key Issues and Debates, 2nd ed.; Bloomsbury Academic: London, UK, 2016; Available online: https://www.bloomsbury.com/uk/education-and-technology-9781350145559/ (accessed on 15 June 2026).
- Rose, D.H.; Meyer, A. Teaching Every Student in the Digital Age: Universal Design for Learning; ASCD: Alexandria, VA, USA, 2002. [Google Scholar]
- CAST. Universal Design for Learning Guidelines Version 2.2 [Graphic Organizer]; Center for Applied Special Technology: Wakefield, MA, USA, 2018; Available online: https://udlguidelines.cast.org (accessed on 15 June 2026).
- Davis, F.D. Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Q. 1989, 13, 319–340. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Venkatesh, V.; Morris, M.G.; Davis, G.B.; Davis, F.D. User acceptance of information technology: Toward a unified view. MIS Q. 2003, 27, 425–478. [Google Scholar] [CrossRef] [Scilit]
- Rei, R.; Stewart, C.; Farinha, A.C.; Lavie, A. COMET: A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 2685–2702. [Google Scholar] [CrossRef] [Scilit]
- Juraska, J.; Finkelstein, M.; Deutsch, D.; Siddhant, A.; Mirzazadeh, M.; Freitag, M. MetricX-23: The Google submission to the WMT 2023 metrics shared task. In Proceedings of the Eighth Conference on Machine Translation (WMT), Singapore, 6–7 December 2023; pp. 1–12. [Google Scholar] [CrossRef] [Scilit]

| RQ | Evidence Used | Main Finding | Conclusion/Interpretation |
|---|---|---|---|
| RQ1 | System implementation and pilot | Offline multilingual broadcasting was feasible in the evaluated university scenarios. | The implemented architecture can support the evaluated offline broadcasting scenarios. |
| RQ2 | Latency analysis and pilot observations | The analytical latency budget estimates 1.8–4.1 s; speech accumulation is the largest assumed component. | The analytical budget and pilot observations indicate that speech accumulation is the main latency contributor. |
| RQ3 | Architecture analysis | Separate language-specific RTP tracks enabled selective audio delivery. | Selective language delivery is supported architecturally. |
| RQ4 | Bandwidth estimation | Estimated outbound bandwidth was approximately 20.76 Mbps for the representative 10-listener scenario. | Bandwidth requirements increase with listener count and require empirical validation. |
| RQ5 | System and deployment analysis | Authentication, recording, transcription, and institutional integration mechanisms were implemented. | The implemented mechanisms support institutional deployment requirements. |
| RQ6 | Institutional analysis and stakeholder feedback | Offline deployment provided potential benefits for data sovereignty, accessibility, connectivity, and institutional control. | These benefits are preliminary and context-dependent rather than statistically generalizable. |
| Pipeline Component | Estimated Latency |
|---|---|
| RMS pre-filter + VAD | 20–50 ms |
| Speech-to-text (STT) | 500–800 ms |
| Translation | 50–200 ms |
| Text-to-speech (TTS) | 100–400 ms |
| AI pipeline total | 670–1450 ms |
| Pipeline Component | Estimated Latency |
|---|---|
| Speech accumulation/segmentation | 1000–2500 ms |
| AI processing | 700–1500 ms |
| RTC delivery (LAN) | 50–100 ms |
| Total expected end-to-end | ≈1.8–4.1 s |
| Component | Per Unit | Total |
|---|---|---|
| Video (1.5 Mbps × 10 listeners) | 1.5 Mbps | 15 Mbps |
| Original audio (64 kbps × 10) | 64 kbps | 0.64 Mbps |
| TTS audio (8 × 64 kbps × 10) | 512 kbps/user | 5.12 Mbps |
| Total estimated outbound | — | ≈20.76 Mbps |
| Benchmark/Metric | Whisper-Large-v3-Turbo (WER) |
|---|---|
| Multilingual average | 7.8–8.4% |
| Male speaker | 8.4% |
| Female speaker | 8.0% |
| Across 99+ languages | ~12% |
| Metric | TranslateGemma 4 B | Note |
|---|---|---|
| MetricX (WMT24++) | 5.32 | Rivals Gemma 3 12 B baseline (4.86); lower is better |
| BLEU | ≈6.95 | Artificially low for LLM output (see below) |
| ChrF | ≈33.28 | More representative of fluency |
| Listeners (Single Broadcast) | Selected-Track (Est.) | Worst-Case All-Track (Est.) |
|---|---|---|
| 10 | ≈15.6 Mbps | ≈20.1 Mbps |
| 50 | ≈78.2 Mbps | ≈100.6 Mbps |
| 100 | ≈156.4 Mbps | ≈201.2 Mbps |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dönmez, E.; Aydin, H. Offline Real-Time Multilingual Broadcasting Using WebRTC and Local AI Services: Architecture, Performance, and Institutional Evaluation. Electronics 2026, 15, 3715. https://doi.org/10.3390/electronics15163715
Dönmez E, Aydin H. Offline Real-Time Multilingual Broadcasting Using WebRTC and Local AI Services: Architecture, Performance, and Institutional Evaluation. Electronics. 2026; 15(16):3715. https://doi.org/10.3390/electronics15163715
Chicago/Turabian StyleDönmez, Erhan, and Hakan Aydin. 2026. "Offline Real-Time Multilingual Broadcasting Using WebRTC and Local AI Services: Architecture, Performance, and Institutional Evaluation" Electronics 15, no. 16: 3715. https://doi.org/10.3390/electronics15163715
APA StyleDönmez, E., & Aydin, H. (2026). Offline Real-Time Multilingual Broadcasting Using WebRTC and Local AI Services: Architecture, Performance, and Institutional Evaluation. Electronics, 15(16), 3715. https://doi.org/10.3390/electronics15163715

