AI-Assisted Vision Alarming System for Blind and Vision- Impaired People
Abstract
1. Introduction
- (1)
- The paper proposes a fit-for-purpose, cost-effective, standalone solution to assist vision-impaired people in daily traffic participation. The system essentially plugs and plays and does not require any Internet connection, edge computing, or online AI processing. To the best of our knowledge, our system is the forefront device from this comprehensive perspective.
- (2)
- The paper develops the first system comprehensively combining all technical components, namely TF Mini, Oak-D Lite, YOLO, TTS, Docker, and Raspberry Pi. Although the use of cameras, TTS or YOLO has been explored in literature, to the best of our knowledge, none of the existing assistive devices has comprehensively combined all of these technical components. Our system can simultaneously perform object detection, obstacle/object distance measurement, and voice feedback generation. The scalability of the system was considered by leveraging Docker, which has not been implemented in many previous works.
- (3)
- The paper proposes an efficient approach for the fusion of the information retrieved from multiple sensing sources with different rates, including the distance information from the TF Mini and distance, depth, and object classification information from the Oak-D Lite in a reliable and optimized way. As shown in the examples later in this paper, VAS can still detect objects accurately even when the Oak-D Lite does not work due to, for instance, the dark conditions or when one sensor (e.g., TF Mini or Oak-D Lite) freezes. This feature plays an important role in the safety of the users in daily traffic participation.
- (4)
- The paper proposes the design of TTS strategies with suitable Quality of Services levels and optimal TTS message structures, carrying the most critical information to provide fast, reliable, but not excessive notifications to the users. To provide sufficient but not overabundant information, the working range is split into critical and non-critical ranges to prioritize the most urgent obstacles. To the best of our knowledge, VAS is the first system that possesses this feature.
- (5)
- The paper develops a novel grouping algorithm where a group of similar and close objects can be notified to the users by just one brief notification, rather than excessive, individual notifications. This is especially helpful in clustered environments. To the best of our knowledge, our system is the first system facilitated by such an algorithm.
2. System Design
2.1. Hardware Selection and Configuration
2.2. Software and Algorithms
2.2.1. ROS 2 Backbone and Architecture
2.2.2. Docker Containerisation and YOLOv4-Tiny Algorithm
2.2.3. Vision Alarming Algorithm
2.3. System Workflow
3. Experiment and Result Analysis
3.1. Experimental Setup
3.2. Results and Analysis
3.3. Discussion
- Real-time Object Detection: VAS utilizes the YOLOv4-tiny model running on the Oak D-Lite AI camera to detect a wide range of objects in real-time scenarios. Previous studies such as [29] reported that electronic travel aids, which VAS belongs to, can help vision impaired users achieve improved object detection, larger detection ranges, and larger safety ranges during obstacle avoidance, in comparison with long canes. Traditional long canes also cannot provide information about objects at higher levels (for example, above waist height) or information about object characteristics, as discussed in [30]. This emphasizes VAS advantages compared to traditional long canes.
- Depth Perception: The Oak D-Lite camera provides stereo depth perception, while the TF Mini adds highly accurate distance readings for objects directly ahead.
- Audio Feedback: A TTS system converts alarming messages with critical information regarding object detection, object distance, and position into easy-to-understand spoken alerts for the user.
- Modular Design: Each sensor or component of VAS operates within its own container, allowing easy upgrades and modifications.
- Standalone Design: VAS does not require the Internet or a GPS connection to operate. The system is designed as a plug-and-play solution; thus, optimizing the convenience of users.
- Cost-effective Design: VAS leverages open-source, cost-effective hardware, middleware, and software, demonstrating its advantage as a relatively low-cost but effective navigation solution.
- Outdoor Navigation: VAS can potentially help users avoid collisions with obstacles when participating in outdoor traffic, with an effective working range of up to 12 m.
- Indoor Navigation: VAS is expected to be useful for avoiding furniture, walls, and other static obstacles when navigating within buildings or indoor environments.
- Low Visibility Navigation: Even in low-light or no-light conditions, VAS still shows the potential to provide consistent, reliable feedback on the presence of obstacles ahead.
- Active Object Localization: VAS may help vision-impaired users actively locate the object or destination of interest like public facilities. In such cases, users can turn the system 360°. If the targeted object is within the working range of VAS, the system can announce to the user about the presence of that object. One example of this scenario is demonstrated in Figure 23a,b where the user can actively locate the bench in the park and along the footpath.
4. Conclusions
Author Contributions
Funding
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AI HAT | Artificial Intelligence Hardware Attached on Top |
| AIoT | Artificial Intelligence of Things |
| BCI | Brain Computer Interface |
| CNN | Convolutional Neural Network |
| COCO | Common Objects in Context |
| CV | Computer Vision |
| DL | Deep Learning |
| FoV | Field of View |
| GPIO | General-Purpose Input/Output |
| GPS | Global Positioning System |
| GSM | Global System for Mobile Communications |
| HDR | High Dynamic Range |
| IoT | Internet of Things |
| LiDAR | Light Detection and Ranging |
| LSTM | Long Short-Term Memory |
| ML | Machine Learning |
| NoIR | No Infrared Filter |
| PubSub | Publisher-Subscriber |
| QoS | Quality of Service |
| RGBD | Red, Green, Blue, and Depth |
| ROS 2 | Robotic Operation System 2 |
| SoX | Sound eXchange |
| SSD | Single Shot Detector |
| ToF | Time of Flight |
| TOPS | Trillion Operations Per Second |
| TTS | Text-To-Speech |
| VAS | Vision Alarming System |
| VGG16 | Visual Geometry Group 16-layer |
| VPU | Vision Processing Unit |
| WHO | World Health Organization |
| YOLO | You Only Look Once |
References
- Vision Australia Website. Available online: https://www.visionaustralia.org/services/eye-conditions/low-vision (accessed on 31 January 2026).
- World Health Organization Website. Available online: https://www.who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment (accessed on 31 January 2026).
- Williams, M.A.; Hurst, A.; Kane, S.K. “Pray before you step out”: Describing personal and situational blind navigation behaviors. In Proceedings of the 15th International ACM SIGACCESS Conference on Computers and Accessibility, Bellevue, WA, USA, 21–23 October 2013; ASSETS ’13. pp. 1–8. [Google Scholar] [CrossRef]
- WeWALK Smart Cane Website. Available online: https://wewalk.io/en/ (accessed on 31 January 2026).
- OrCam MyEye Camera Website. Available online: https://www.quantumrlv.com.au/products/orcam-myeye-pro-text-to-speech-wearable?srsltid=AfmBOoo2I-dRo8C4HaEzd02sYsFN_-zZpfaidBk6K8xmhUPtOCJk-igV (accessed on 31 January 2026).
- Poggi, M.; Mattoccia, S. A wearable mobility aid for the visually impaired based on embedded 3D vision and deep learning. In Proceedings of the 2016 IEEE Symposium on Computers and Communication (ISCC), Messina, Italy, 27–30 June 2016; pp. 208–213. [Google Scholar] [CrossRef]
- Safiya, K.M.; Pandian, R. Real-Time Photo Captioning for Assisting Blind and Visually Impaired People Using LSTM Framework. IEEE Sens. Lett. 2023, 7, 6008804. [Google Scholar] [CrossRef]
- Mohanraj, P.; Rajasekar, T.; Sivaelango, N.; Sri Karthickraja, V.; Vignesh, N. Wearable Device for Visually Impaired using Deep Learning. In Proceedings of the 2024 3rd International Conference on Applied Artificial Intelligence and Computing (ICAAIC), Salem, India, 5–7 June 2024; pp. 348–352. [Google Scholar] [CrossRef]
- Masud, U.; Saeed, T.; Malaikah, H.M.; Islam, F.U.; Abbas, G. Smart Assistive System for Visually Impaired People Obstruction Avoidance Through Object Detection and Classification. IEEE Access 2022, 10, 13428–13441. [Google Scholar] [CrossRef]
- Rahman, M.A.; Sadi, M.S. IoT Enabled Automated Object Recognition for the Visually Impaired. Comput. Methods Programs Biomed. Update 2021, 1, 100015. [Google Scholar] [CrossRef]
- Jivrajani, K.; Patel, S.K.; Parmar, C.; Surve, J.; Ahmed, K.; Bui, F.M.; Al-Zahrani, F.A. AIoT-Based Smart Stick for Visually Impaired Person. IEEE Trans. Instrum. Meas. 2023, 72, 2501311. [Google Scholar] [CrossRef]
- Krishna, A.N.; Chaitra, Y.L.; Bharadwaj, A.M.; Abbas, K.T.; Abraham, A.; Prasad, A.S. Real-time Machine Vision System for the Visually Impaired. SN Comput. Sci. 2024, 5, 399. [Google Scholar] [CrossRef]
- Koteswararao, M.; Karthikeyan, P.R. Comparative Analysis of YOLOv3-320 and YOLOv3-tiny for the Optimised Real-Time Object Detection System. In Proceedings of the 2022 3rd International Conference on Intelligent Engineering and Management (ICIEM), London, UK, 27–29 April 2022; pp. 495–500. [Google Scholar] [CrossRef]
- Madake, J.; Bhatlawande, S.; Mahale, S.; Mahankaliwar, S. VisioGuide: Modified YoloV7 with ResNext for Object Detection to Aid Visually Impaired Individuals. In Proceedings of the 2024 International Conference on Current Trends in Advanced Computing (ICCTAC), Bengaluru, India, 8–9 May 2024; pp. 1–9. [Google Scholar] [CrossRef]
- Panwar, M.; Rajoria, H.; Soreng, J.; Batsyas, R.; Kharangarh, P.R. Review—Innovations in Flexible Sensory Devices for the Visually Impaired. ECS J. Solid State Sci. Technol. 2024, 13, 077011. [Google Scholar] [CrossRef]
- Raspberry Pi 5 Website. Available online: https://www.raspberrypi.com/products/raspberry-pi-5/ (accessed on 31 January 2026).
- Oak-D Lite Camera Website. Available online: https://shop.luxonis.com/products/oak-d-lite-1?variant=42583102456031 (accessed on 31 January 2026).
- TF Mini Website. Available online: https://au.mouser.com/datasheet/3/1361/1/SJ_GU_TFmini_Plus_01_A03_Datasheet_EN.pdf (accessed on 31 January 2026).
- TF Mini Datasheet Website. Available online: https://www.mouser.com/datasheet/2/1099/Benewake_10152020_TFmini_Plus-1954028.pdf (accessed on 31 January 2026).
- Liu, X.; Zhang, S.; Zeng, J.; Fan, F. Analysis and optimization strategy of travel system for urban visually impaired people. Sustainability 2019, 11, 1735. [Google Scholar] [CrossRef]
- ROS 2 Website. Available online: https://www.ros.org/ (accessed on 31 January 2026).
- ROS 2 Publisher for Oak-D-Lite Website. Available online: https://www.sundance.com/oak-d-lite/ (accessed on 31 January 2026).
- Piper Guideline Website. Available online: https://noerguerra.com/how-to-read-text-aloud-with-piper-and-python/ (accessed on 31 January 2026).
- Bittner, R.; Humphrey, E.; Bello, J. Pysox: Leveraging the Audio Signal Processing Power of SoX in Python. In Proceedings of the International Society for Music Information Retrieval Conference: Late-Breaking and Demo Papers, Brooklyn, NY, USA, 11 August 2016. [Google Scholar]
- Docker Website. Available online: https://docs.docker.com/get-started/docker-overview/ (accessed on 31 January 2026).
- Wang, L.; Zhou, K.; Chu, A.; Wang, G.; Wang, L. An improved light-weight traffic sign recognition algorithm based on YOLOv4-tiny. IEEE Access 2021, 9, 124963–124971. [Google Scholar] [CrossRef]
- Kahaki, Z.R.; Safarpour, A.R.; Daneshmandi, H. The spatiotemporal gait parameters among people with visual impairment: A literature review study. Oman J. Ophthalmol. 2023, 16, 427–433. [Google Scholar] [CrossRef] [PubMed]
- System Demonstration. Available online: https://youtu.be/Od-RasPf474 (accessed on 4 April 2026).
- Jin, R.; Petoe, M.A.; McCarthy, C.D.; Serra, J.R.; Starkey, S.; McGinley, J.; Ayton, L.N. Functional performance comparison of long cane and secondary electronic travel aids for mobility enhancement. Br. J. Vis. Impair. 2024, 02646196241285098. [Google Scholar] [CrossRef]
- Pittet, C.E.; Ortega, E.V.; Fabien, M.; Wallace, M.T.; Gori, M.; Murray, M.M. Efficacy of electronic travel aids for the blind and visually impaired during wayfinding. Sci. Rep. 2026, 16, 6423. [Google Scholar] [CrossRef] [PubMed]
- Phung, S.L.; Nguyen, T.N.A.; Le, H.T.; Chapple, P.B.; Ritz, C.H.; Bouzerdoum, A.; Tran, L.C. Mine-like object sensing in sonar imagery with a compact deep learning architecture for scarce data. In Proceedings of the 2019 Digital Image Computing: Techniques and Applications (DICTA), Perth, WA, Australia, 2–4 December 2019; pp. 1–7. [Google Scholar] [CrossRef]
- Wisanmongkol, J.; Taparugssanagorn, A.; Tran, L.C.; Le, A.T.; Huang, X.; Ritz, C.; Dutkiewicz, E.; Phung, S.L. An ensemble approach to deep-learning-based wireless indoor localization. IET Wirel. Sens. Syst. 2022, 12, 33–55. [Google Scholar] [CrossRef]
- Alshbatat, A.I.N.; Vial, P.J.; Premaratne, P.; Tran, L.C. EEG-based Brain-computer Interface for Automating Home Appliances. J. Comput. 2014, 9, 2159–2166. [Google Scholar]























| Paper | Reduction of User Training | Lightweight/ Compactness | GPS Requirement | Sound/ Tactile Feedback | Internet Connection Independence | Objects Variety |
|---|---|---|---|---|---|---|
| [12] | Yes | Yes | Partial | Yes | No | No |
| [11] | Partial | Yes | Partial | Yes | Partial | Yes |
| [10] | Partial | Partial | Partial | Yes | Partial | Partial |
| [9] | Partial | Yes | No | Yes | Yes | Yes |
| [6] | Partial | Partial | Partial | Yes | Yes | Partial |
| This paper | Yes | Yes | No | Yes | Yes | Yes |
| Description | Parameter Value |
|---|---|
| Operating range | 0.1–12 m |
| Accuracy | ±5 cm (0.1–6 m) |
| ±1% (6–12 m) | |
| Measurement unit | cm |
| Distance resolution | 1 cm |
| Field of View (FoV) | 3.6° |
| Frame rate | 1–1000 Hz (adjustable) |
| Scenario | Precision | Reliability |
|---|---|---|
| Indoor office (sufficient light) | 88.8% | High |
| Indoor office (insufficient light) | 75.6% | High |
| Indoor office (no light) | N/A | High |
| Outdoor (sunny) | 90.8% | High |
| Outdoor (sunny-blurry) | 77.5% | Acceptable |
| Outdoor (night) | 77% | High |
| Metric | Value |
|---|---|
| Processing latency | 100 ms |
| Power consumption | 12–15 W |
| Docker usage | 35–55% CPU |
| Highest localization accuracy | 5 cm |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tran, L.C.; Ly, S.K.; Blacklidge, R.; Shemmell, J.; Difford, N.; Cox, D.E.; Harada, T. AI-Assisted Vision Alarming System for Blind and Vision- Impaired People. Sensors 2026, 26, 2929. https://doi.org/10.3390/s26102929
Tran LC, Ly SK, Blacklidge R, Shemmell J, Difford N, Cox DE, Harada T. AI-Assisted Vision Alarming System for Blind and Vision- Impaired People. Sensors. 2026; 26(10):2929. https://doi.org/10.3390/s26102929
Chicago/Turabian StyleTran, Le Chung, Sinh Khai Ly, Rhys Blacklidge, Jonathan Shemmell, Nathan Difford, Daniel Edward Cox, and Theresa Harada. 2026. "AI-Assisted Vision Alarming System for Blind and Vision- Impaired People" Sensors 26, no. 10: 2929. https://doi.org/10.3390/s26102929
APA StyleTran, L. C., Ly, S. K., Blacklidge, R., Shemmell, J., Difford, N., Cox, D. E., & Harada, T. (2026). AI-Assisted Vision Alarming System for Blind and Vision- Impaired People. Sensors, 26(10), 2929. https://doi.org/10.3390/s26102929

