Privacy-Preserving Smart Glasses Navigation Through VLM Fine-Tuning for People with Visual Impairments
Abstract
1. Introduction
- An in-depth requirements analysis based on expert interviews and a user survey involving visually impaired individuals.
- A LoRA-based fine-tuning approach that incorporates target-group-specific knowledge and runs locally on a smartphone.
- An evaluation comparing image descriptions generated by LLMs with descriptions of the same images provided by domain experts.
2. Related Work
3. Approach
3.1. Hardware Design Considerations
- 1.
- Inference on cloud server.Because inference is performed on cloud servers, data is transmitted to external data centers for processing. This enables the use of high-performance, up-to-date models without the need for dedicated hardware. On the other hand, however, this raises significant challenges regarding data protection and data security, as sensitive information (belonging to the user and possibly third parties) is transmitted and potentially processed outside the user’s direct control. Furthermore, availability is heavily dependent on a stable internet connection. Local resource requirements, however, are low, as processing takes place entirely on external infrastructure. In contrast, this generates ongoing operating costs, which can vary depending on usage and model complexity and may be passed on to the user.
- 2.
- Inference on private server.Performing the inference on own servers provides greater control over the data and processing workflows compared to cloud solutions. This allows sensitive data to be processed within a controlled infrastructure, which offers advantages in terms of data protection and security. However, potential risks related to data transmission may still exist. The performance of the models on own servers is generally lower than that of major cloud providers, but is still very high and certainly sufficient for our application. This option requires more resources, as it involves operating and maintaining your own hardware. Additionally, there are ongoing operating costs for administration, maintenance, and infrastructure. Availability also depends on a stable network connection.
- 3.
- Inference on smart phone.When performing the inference on smart phones, data processing takes place directly on the device itself, so the data does not need to leave the device in this case. This offers significant advantages in terms of data privacy and security, as well as high availability, since there is no reliance on external infrastructure. Besides reduced inference speed, the local processing results in only very low latency. However, significant limitations arise from the limited computing power of mobile devices. Therefore, only simplified models can be used. Resource requirements are constrained by energy consumption, limited battery capacity, and thermal limitations. Operating costs, on the other hand, are low, as no external infrastructure is required and no recurring fees are incurred.
- 4.
- Inference on smart glasses.The inference on smart glasses represents an even more sophisticated implementation of edge processing, in which data is processed directly on wearable devices. Similar to the option of inference on a smartphone, this means that the data remains on the device, which offers advantages for data protection and security. In addition, the low system complexity enables high availability without dependence on external infrastructure and as it improves system reliability. However, hardware capacities and implementation options are currently very limited. This allows, if at all, only the use of drastically reduced models. In many cases, fully local inference is not yet practical. The available resources of the smart glasses are still very limited, particularly in terms of power supply and computing power. Operating costs are currently difficult to quantify. In comparison, however, these are on the lower end of the spectrum, since while powerful smart glasses are necessary, there are no ongoing usage fees, as is the case with the smartphone option.
3.2. User-Specific Design of VLM
3.2.1. Need for a Specialized VLM
3.2.2. Need for LLM Fine-Tuning
3.2.3. Existing Training Data
3.3. User-Centered Requirements Analysis
3.3.1. Online Survey
- Methods
- Participants
- Apparatus and Materials
- Procedure
- Results
3.3.2. Expert Interview
- Methods
- Participants
- Apparatus and Materials
- Procedure
- Results
3.3.3. Discussion
3.4. Design Considerations for LLM Answer
4. Implementation
4.1. System Architecture
4.2. VLM Architecture and Training
4.2.1. Data Source and Preparation
4.2.2. Prompt Engineering
). This wording is based on the statements of the mobility trainers interviewed, who emphasize navigation-related information. One expert explains: “Only obstacles that are in the way should be mentioned.” Another adds: “It’s enough to know where I am and whether it’s a pedestrian zone throughout.” The statement “The two passersby… are irrelevant when I walk past on the left side” also illustrates that information is primarily evaluated based on its impact on safe movement. The prompt instruction to focus on mobility rather than a general scene description was therefore derived directly from the experts’ statements.
). In the expert interviews as well, it repeatedly came up that dangerous obstacles should take precedence over other information. As one mobility trainer explains: “If a dangerous obstacle comes up, that’s what needs to be mentioned first.” He further elaborates: “Above all, precipices such as stairs must be identified at the very beginning.” In addition, it is emphasized: “Level changes in general, and anything sticking out, such as branches, must be identified at the very beginning.” The order specified in the prompt (see Appendix A (
) is thus based on the safety relevance of the obstacles as described by the experts.
). The relevance of such obstacles is supported by both the expert interviews and the quantitative study. In the survey, obstacles at head height received the second-highest level of agreement among the types of objects evaluated. The expert interviews also highlight the importance of such obstacles, stating: “Branches must be identified first.” This could be related to the fact that ground obstacles are often detected by the long cane, whereas obstacles at head or chest height are more easily overlooked.
). Several interviewees emphasize the importance of such guide lines. This is reflected in the statement: “There is a guide line on the left.” Another recommends: “Follow the guide line on the way to the left, that is, the house wall or the bridge wall.” He also pointed out: “I recommend that the blind person feel their way close to the house wall and use this wall as a tactile guide line.” These guidance aids thus largely correspond to the recommendations of mobility trainers.
) A recurring theme in the interviews is that people do not constitute relevant obstacles in most situations. This assessment is also reflected in the statement: “The description is quite good, except for the people on the right who are walking off. Another adds: “The two people on the right are walking away from you and are not obstacles. The people coming toward you are also not obstacles.” A third expert makes a similar point: “(…) but I don’t describe people, because they’re somewhere else entirely the next second.” This illustrates why people usually don’t need to be mentioned.4.2.3. Training LoRa
5. Results
5.1. Exemplary Results
5.2. Evaluation and Discussion
5.2.1. Qualitative Evaluation
5.2.2. Quantitative Evaluation
5.2.3. Evaluation Key Outcomes
- A thorough requirements analysis (involving the target user group and experts) is essential for identifying elements that are relevant to blind people.
- It is not only the details in the image that are relevant to blind people that matter, but also the use of appropriate specialist terminology.
- Prompt engineering is beneficial, but its effectiveness is limited on its own. An adapted model with enriched knowledge for blind people has proven to be significantly more effective.
- Larger models (on servers) do deliver better results, but they are also limited by a lack of knowledge about the target group.
6. Conclusions
Limitations and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. System Prompt Used for Approach
,
,
) serve as references indicating which parts of the prompt were created based on the requirements that were developed.- Your job is to give one natural spoken instruction for the next part of walking. Focus on safe mobility, not scene description.
- Prioritize in this order:
- Immediate danger.
- Obstacles a white cane may not reliably detect, especially above knee height.
- Chest-height, head-height, overhanging, objects sticking out and protruding obstacles.
- Ground obstacles in the walking path.
- Whether a passable pedestrian corridor exists for normal cane sweep and body passage.
- Cyclists or other traffic conflicting with the pedestrian path.
- Crossings, tactile paving, curb ramps, curb edges, curb alignment, steps, ramps, stairs, signal poles, and traffic islands.
- One or two useful guidance cues for the next movement.
Hazard ranking:- -
- Prioritize the most immediate safety-critical hazard for the next movement, even if another object is more visually prominent.
- -
- Do not focus on general landmarks
- -
- Instead focus on objects, obstacles and landmarks that are helpful for a blind white cane user such as: bollards, poles, or curb edges when barrier tape, construction barriers, protruding objects, overhanging hazards, vegetation sticking out or blocked pedestrian corridors are more important.
- -
- If a pedestrian barrier, tape, or blocked corridor affects the next movement, mention it before guidance cues.
- Core safety rules:
- -
- Warn clearly about hazards not reliably detectable by normal white cane sweep before body contact.
- -
- Give special attention to head-height, chest-height, overhanging, and protruding obstacles. Always mention them.
- -
- This includes signs, poles, mirrors, awnings, balconies, café umbrellas, scaffolding, construction sites, window boxes, building fixtures, projecting shelves, trailer hitches, handlebars, metal bars, and other objects protruding from façades, parked vehicles, or street furniture.
- -
- If upper-body risk is present, you may say: protect your face with your free hand, or keep your free hand at head height.
- -
- If you give a protective action such as protecting the face or keeping the free hand at head height or chest height, briefly name the protruding, overhanging, or upper-body hazard that makes it necessary.
- -
- Do not give protective-hand instructions without naming the nearby hazard they are meant to prevent.
- -
- Do not call the path blocked if a short safe pedestrian corridor is still visible.
- -
- If the path is narrow but passable, say continue carefully, pass slightly left, or pass slightly right.
- -
- Treat people and temporary crowding as movable unless no safe bypass is visible.
- -
- If no clearly walkable short-horizon route is visible, say stop and choose another way, or stop and ask someone nearby for help.
- -
- If safety is uncertain, give the safer instruction.
- Guidance cues:
- -
- Prefer immediate useful cues such as tactile paving, guiding line, curb edge, curb ramp, curb alignment, signal pole, bollard, handrail, traffic island, railing, wall line, fence line, bridge wall, building edge, house wall, or continuous façade.
- -
- Prefer cues ahead over distant side boundaries.
- -
- Recommend a side guiding line only if it is near, reachable, and usable.
- -
- If a wall line, railing, curb edge, or façade is cluttered by bikes, signs, planters, café furniture, scooters, barriers, trailer hitches, or other protrusions, do not present it as a clean guide without warning.
- -
- In open areas, prefer forward cues such as tactile paving, a bollard, signal pole, curb ramp, curb alignment, or the clear pedestrian corridor.
- - Treat people and temporary crowding as movable unless no safe bypass is visible.
- -
- Do not describe the entire scene Direction and distance:
- -
- Use egocentric directions only.
- -
- Use left/right wording for movement and clock-face directions for nearby hazards or cues when helpful.
- -
- Keep distances approximate: steps under 10 m, rounded meters from 10 to 20 m, and only coarse bands or landmarks beyond 20 m.
- -
- Never claim precise measurements.
- Style:
- -
- Use mobility-oriented spoken language.
- -
- Prefer these terms when accurate: white cane, cane sweep, pedestrian corridor, passable corridor, tactile paving, guiding line, curb edge, curb ramp, curb alignment, signal pole, traffic island, head-height obstacle, chest-height obstacle, overhanging obstacle, protruding obstacle.
- -
- Do not describe the whole scene.
- -
- Do not explain reasoning.
- -
- Do not output labels, bullets, JSON, or multiple options.
References
- Bourne, R.R.A.; Flaxman, S.R.; Braithwaite, T.; Cicinelli, M.V.; Das, A.; Jonas, J.B.; Keeffe, J.; Kempen, J.H.; Leasher, J.; Limburg, H.; et al. Magnitude, temporal trends, and projections of the global prevalence of blindness and distance and near vision impairment: A systematic review and meta-analysis. Lancet Glob. Health 2017, 5, e888–e897. [Google Scholar] [CrossRef] [PubMed]
- World Report on Vision; Technical Report; World Health Organisation: Geneva, Switzerland, 2019.
- Naayini, P.; Myakala, P.K.; Bura, C.; Jonnalagadda, A.K.; Kamatala, S. AI-Powered Assistive Technologies for Visual Impairment. arXiv 2025, arXiv:2503.15494. [Google Scholar] [CrossRef]
- Udayakumar, D.; Gopalakrishnan, S.; Raghuram, A.; Kartha, A.; Krishnan, A.K.; Ramamirtham, R.; Muthangi, R.; Raju, R. Artificial intelligence-powered smart vision glasses for the visually impaired. Indian J. Ophthalmol. 2025, 73, S492–S497. [Google Scholar] [CrossRef] [PubMed]
- Danish, S.; Sadeghi-Niaraki, A.; Khan, S.U.; Dang, L.M.; Tightiz, L.; Moon, H. A comprehensive survey of Vision–Language Models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets. Inf. Fusion 2026, 126, 103623. [Google Scholar] [CrossRef]
- Shou, Z.; Lin, F. Enhancing Semantic Understanding in Vision Language Models Using Meaning Representation Negative Generation. In Proceedings of the Fourth Workshop on Knowledge-Infused Learning, Barcelona, Spain, 25–26 August 2024. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Computer Vision—ECCV 2014; Series Title: Lecture Notes in Computer Science; Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T., Eds.; Springer International Publishing: Cham, Switzerland, 2014; Volume 8693, pp. 740–755. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the 10th International Conference on Learning Representations (ICLR 2022), Online, 25–29 April 2022. [Google Scholar]
- Meijer, P. An experimental system for auditory image representations. IEEE Trans. Biomed. Eng. 1992, 39, 112–121. [Google Scholar] [CrossRef] [PubMed]
- Aymaz, Ş.; Çavdar, T. Ultrasonic Assistive Headset for visually impaired people. In Proceedings of the 2016 39th International Conference on Telecommunications and Signal Processing (TSP); IEEE: Piscataway, NJ, USA, 2016; pp. 388–391. [Google Scholar] [CrossRef]
- Busaeed, S.; Katib, I.; Albeshri, A.; Corchado, J.M.; Yigitcanlar, T.; Mehmood, R. LidSonic V2.0: A LiDAR and Deep-Learning-Based Green Assistive Edge Device to Enhance Mobility for the Visually Impaired. Sensors 2022, 22, 7435. [Google Scholar] [CrossRef] [PubMed]
- Kiuru, T.; Metso, M.; Utriainen, M.; Metsävainio, K.; Jauhonen, H.M.; Rajala, R.; Savenius, R.; Ström, M.; Jylhä, T.N.; Juntunen, R.; et al. Assistive device for orientation and mobility of the visually impaired based on millimeter wave radar technology—Clinical investigation results. Cogent Eng. 2018, 5, 1450322. [Google Scholar] [CrossRef]
- Kreimeier, J.; Kappe, M.; Götzelmann, T. BlindScanLine: Preliminary Implementation and Evaluation of a Cross-Platform Line Scanning Sonification Approach Comparing Frequency and Amplitude Modulation. In Proceedings of the 13th ACM International Conference on PErvasive Technologies Related to Assistive Environments; ACM: New York, NY, USA, 2020; PETRA ’20. [Google Scholar] [CrossRef]
- Bułat, J.; Głowacz, A. Vision-based navigation assistance for visually impaired individuals using general purpose mobile devices. In Proceedings of the 2016 International Conference on Signals and Electronic Systems (ICSES); IEEE: Piscataway, NJ, USA, 2016; pp. 189–194. [Google Scholar] [CrossRef]
- Nguyen, M.; Le, H.; Yan, W.Q.; Dawda, A. A Vision Aid for the Visually Impaired using Commodity Dual-Rear-Camera Smartphones. In Proceedings of the 2018 25th International Conference on Mechatronics and Machine Vision in Practice (M2VIP), Stuttgart, Germany, 20–22 November 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 1–6. [Google Scholar] [CrossRef]
- Nancy, V.; Balakrishnan, G. Thermal Image-Based Object Classification for Guiding the Visually Impaired. Comput. J. 2021, 64, 1747–1759. [Google Scholar] [CrossRef]
- Lee, C.-H.; Su, Y.-C.; Chen, L.-G. An intelligent depth-based obstacle detection system for visually-impaired aid applications. In Proceedings of the 2012 13th International Workshop on Image Analysis for Multimedia Interactive Services, Dublin, Ireland, 23–25 May 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 1–4. [Google Scholar] [CrossRef]
- Lee, Y.H.; Medioni, G. RGB-D camera based wearable navigation system for the visually impaired. Comput. Vis. Image Underst. 2016, 149, 3–20. [Google Scholar] [CrossRef]
- Kreimeier, J.; Götzelmann, T. Real World VR Proxies to Support Blind People in Mobility Training. In Mensch und Computer 2018-Workshopband; Gesellschaft für Informatik e.V.: Bonn, Germany, 2018. [Google Scholar] [CrossRef]
- Mukhiddinov, M.; Cho, J. Smart Glass System Using Deep Learning for the Blind and Visually Impaired. Electronics 2021, 10, 2756. [Google Scholar] [CrossRef]
- Younis, O.; Al-Nuaimy, W.; Alomari, M.H.; Rowe, F. A Hazard Detection and Tracking System for People with Peripheral Vision Loss using Smart Glasses and Augmented Reality. Int. J. Adv. Comput. Sci. Appl. 2019, 10, 1–9. [Google Scholar] [CrossRef]
- Katzschmann, R.K.; Araki, B.; Rus, D. Safe Local Navigation for Visually Impaired Users With a Time-of-Flight and Haptic Feedback Device. IEEE Trans. Neural Syst. Rehabil. Eng. 2018, 26, 583–593. [Google Scholar] [CrossRef] [PubMed]
- Erdaw, H.B.; Taye, Y.G.; Lemma, D.T. A Real-Time Obstacle Detection and Classification System For Assisting Blind and Visually Impaired People Based On Yolo Model. In Proceedings of the 2023 International Conference on Information and Communication Technology for Development for Africa (ICT4DA); IEEE: Piscataway, NJ, USA, 2023; pp. 79–84. [Google Scholar] [CrossRef]
- Tahoun, N.; Awad, A.; Bonny, T. Smart Assistant for Blind and Visually Impaired People. In Proceedings of the 2019 3rd International Conference on Advances in Artificial Intelligence, Istanbul, Turkey, 26–28 October 2019; pp. 227–231. [Google Scholar] [CrossRef]
- Yee, L.R.; Kamaludin, H.; Safar, N.Z.M.; Wahid, N.; Abdullah, N.; Meidelfi, D. Intelligence Eye for Blinds and Visually Impaired by Using Region-Based Convolutional Neural Network (R-CNN). JOIV Int. J. Inform. Vis. 2021, 5, 409. [Google Scholar] [CrossRef]
- Ikram, S.; Sarwar Bajwa, I.; Gyawali, S.; Ikram, A.; Alsubaie, N. Enhancing Object Detection in Assistive Technology for the Visually Impaired: A DETR-Based Approach. IEEE Access 2025, 13, 71647–71661. [Google Scholar] [CrossRef]
- Kamikubo, R.; Kayukawa, S.; Kaniwa, Y.; Wang, A.; Kacorri, H.; Takagi, H.; Asakawa, C. Beyond Omakase: Designing Shared Control for Navigation Robots with Blind People. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan, 26 April–1 May 2025; pp. 1–17. [Google Scholar] [CrossRef] [PubMed]
- Oumard, C.; Kreimeier, J.; Götzelmann, T. Implementation and Evaluation of a Voice User Interface with Offline Speech Processing for People who are Blind or Visually Impaired. In Proceedings of the 15th PErvasive Technologies Related to Assistive Environments Conference; ACM Press: New York, NY, USA, 2022; PETRA 2021; 9p. [Google Scholar] [CrossRef]
- Kaniwa, Y.; Kuribayashi, M.; Kayukawa, S.; Sato, D.; Takagi, H.; Asakawa, C.; Morishima, S. ChitChatGuide: Conversational Interaction Using Large Language Models for Assisting People with Visual Impairments to Explore a Shopping Mall. Proc. ACM Hum.-Comput. Interact. 2024, 8, 1–25. [Google Scholar] [CrossRef]
- Zhou, G.; Hong, Y.; Wu, Q. NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models. arXiv 2023, arXiv:2305.16986. [Google Scholar] [CrossRef]
- Seo, J.; Kamath, S.S.; Zeidieh, A.; Venkatesh, S.; McCurry, S. MAIDR Meets AI: Exploring Multimodal LLM-Based Data Visualization Interpretation by and with Blind and Low-Vision Users. In Proceedings of the The 26th International ACM SIGACCESS Conference on Computers and Accessibility, St. John’s, NL, Canada, 27–30 October 2024; pp. 1–31. [Google Scholar] [CrossRef]
- AI Smart Glasses for Accessibility|Ally Solos & Envision. 2025. Available online: https://www.ally.me/glasses/solos (accessed on 17 December 2025).
- Be My AI. 2025. Available online: https://www.bemyeyes.com/bme-ai/ (accessed on 9 January 2026).
- Envision—Assistive Technology for Blind and Low Vision. 2025. Available online: https://www.letsenvision.com/ (accessed on 17 December 2025).
- Qi, X.; Panda, A.; Lyu, K.; Ma, X.; Roy, S.; Beirami, A.; Mittal, P.; Henderson, P. Safety Alignment should be Made More Than Just a Few Tokens Deep. arXiv 2024, arXiv:2406.05946. [Google Scholar]
- Li, X.; Gandhi, B.; Zhan, M.; Nehra, M.; Zhang, Z.; Sun, Y.; Song, M.; Zhang, N.; Wang, X. Fine-Tuning Vision-Language Models for Visual Navigation Assistance. arXiv 2025, arXiv:2509.07488. [Google Scholar] [CrossRef]
- Chao, A.; Maquiling, E.; Chao, E.; Sanjeev, R.; Bossen, T.; Greer, R. Automated Context-Aware Navigation Support for Individuals with Visual Impairment Using Multimodal Language Models in Urban Environments. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE Computer Society: Washington, DC, USA, 2025; pp. 6745–6752. [Google Scholar] [CrossRef]
- Yuan, Z.; Zhang, T.; Zhu, Y.; Zhang, J.; Deng, Y.; Jia, Z.; Luo, P.; Duan, X.; Zhou, J.; Zhang, J. WalkVLM: Aid Visually Impaired People Walking by Vision Language Model. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE Computer Society: Washington, DC, USA, 2025; pp. 9845–9854. [Google Scholar] [CrossRef]
- Yang, B.; He, L.; Liu, K.; Yan, Z. VIAssist: Adapting Multi-Modal Large Language Models for Users with Visual Impairments. In Proceedings of the 2024 IEEE International Workshop on Foundation Models for Cyber-Physical Systems & Internet of Things (FMSys), Hong Kong, China, 13–15 May 2024; pp. 32–37. [Google Scholar] [CrossRef]
- Tokmurziyev, I.; Cabrera, M.A.; Khan, M.H.; Mahmoud, Y.; Tsetserukou, D. LLM-Glasses: GenAI-driven Glasses with Haptic Feedback for Navigation of Visually Impaired People. arXiv 2026, arXiv:2503.16475. [Google Scholar] [CrossRef]
- Liu, X.; Lee, D.; Gonzalez, E.J.; Gonzalez-Franco, M.; Suzuki, R. VisionClaw: Always-On AI Agents through Smart Glasses. arXiv 2026, arXiv:2604.03486. [Google Scholar]
- Sun, Z. A Short Survey of Viewing Large Language Models in Legal Aspect. arXiv 2023, arXiv:2303.09136. [Google Scholar]
- Xu, N.; Zhang, J.; Li, C.; An, H.; Zhou, C.; Wang, J.; Xu, B.; Li, Y.; Du, T.; Ji, S. Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content? arXiv 2025, arXiv:2512.21871. [Google Scholar]
- Li, X.L.; Liang, P. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Online, 1–6 August 2021; pp. 4582–4597. [Google Scholar] [CrossRef]
- Sharma, P.; Ding, N.; Goodman, S.; Soricut, R. Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Melbourne, Australia, 15–20 July 2018; pp. 2556–2565. [Google Scholar] [CrossRef]
- Wu, W.; Zhao, Y.; Chen, H.; Gu, Y.; Zhao, R.; He, Y.; Zhou, H.; Shou, M.Z.; Shen, C. DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models. arXiv 2023, arXiv:2308.06160. [Google Scholar] [CrossRef]
- Srinivasan, K.; Raman, K.; Chen, J.; Bendersky, M.; Najork, M. WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, 11–15 July 2021; pp. 2443–2449. [Google Scholar] [CrossRef]
- Gurari, D.; Li, Q.; Stangl, A.J.; Guo, A.; Lin, C.; Grauman, K.; Luo, J.; Bigham, J.P. VizWiz Grand Challenge: Answering Visual Questions from Blind People. arXiv 2018, arXiv:1802.08218. [Google Scholar] [CrossRef]
- Chintalapati, S.; Bragg, J.; Wang, L.L. A Dataset of Alt Texts from HCI Publications: Analyses and Uses Towards Producing More Descriptive Alt Texts of Data Visualizations in Scientific Papers. In Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility; ACM: New York, NY, USA, 2022; pp. 1–12. [Google Scholar] [CrossRef]
- Jiang, Z.; Yuan, X.; Qu, H.; Lin, S.; Liu, K.; Fan, W.; Li, Q. SuperGlasses: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glasses. arXiv 2026, arXiv:2602.22683. [Google Scholar]
- SoSci Survey. Befragung Sehbehinderter Personen. 2015. Available online: https://www.soscisurvey.de/help/doku.php/de:create:barriers_visual (accessed on 24 May 2026).
- Islam, M.T.; Kabir, I.; Pearce, E.A.; Reza, M.A.; Billah, S.M. Identifying Crucial Objects in Blind and Low-Vision Individuals’ Navigation. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility; ACM: New York, NY, USA, 2024; ASSETS ’24. [Google Scholar] [CrossRef]
- Gemma 4 E2B. 2026. Available online: https://huggingface.co/google/gemma-4-E2B-it/ (accessed on 15 April 2026).
- Gemma 4 31B. 2026. Available online: https://huggingface.co/google/gemma-4-31B/ (accessed on 15 April 2026).








| Data Protection | Availability | Performance | Resource Requirements | OperationalCosts | |
|---|---|---|---|---|---|
| Cloud Server | −− | − | ++ | + | − |
| Private Server | + | − | + | − | − |
| Smart Phone | ++ | ++ | +/− | − | + |
| Smart Glasses | ++ | ++ | − | −− | +/− |
![]() | |
|---|---|
| Expert 1 | Ahead of you is a construction site next to the building, surrounded by barrier tape. Slow down as you approach the barrier tape, and walk along it until you reach the end of the construction site. There, the barrier tape is laid out toward the building wall. Walk along it. |
| Expert 2 | On the sidewalk ahead of you, there is a barrier on the left, inner edge, approximately 10 m long, consisting of barrier tape and three bollards. There is a tree on the right; the path is uneven but passable. |
| Expert 3 | Ahead of you is the wide sidewalk on the left side of the street. At the inner guiding line, there is a temporary construction barrier that you must go around in the middle of the sidewalk. At the outer guiding line, there are tree guards spaced far apart. |
| Gemma4 31B | Stop. Barrier tape and bollards block the path ahead. Move right to go around them, but be mindful of a tree support structure a few steps ahead at 2 o’clock. Use the sidewalk edge on your right for guidance. |
| Gemma4 E2B | Continue straight ahead along the sidewalk; watch for the wooden structure on your right. |
| Gemma4 E2B + LoRa | Continue straight ahead, watch for the red and white barrier tape across the path, keep to the right of the tape. Use the curb edge on your right as a guide. |
![]() | |
|---|---|
| Expert 1 | The sidewalk is blocked by scaffolding on your right against the building wall. Feel your way further to the left toward the sloping curb, and walk along it for 10 m. Then you will have passed the construction site. |
| Expert 2 | There is scaffolding on the right along the inner guide line. Continue walking right along the outer guiding line (curb). After about 10 m, the construction site ends. Pay attention to the exact width of your cane swing. The sidewalk narrows to less than 1 m. |
| Expert 3 | The inner guiding line on the sidewalk ahead of you is blocked by several obstacles; follow the sidewalk and use the outer guiding line for orientation. It is unclear if the inner guiding line will be usable again. |
| Gemma4 31B | Continue forward carefully, passing slightly left. Protect your face with your free hand due to protruding scaffolding poles on your right. Use the curb edge on your left as a guide. |
| Gemma4 E2B | Continue along the paved path, watch for the curb edge on your right. |
| Gemma4 E2B + LoRa | Continue straight, watch for the curb edge on your right, and be aware of the large metal scaffolding structure on the right side of the path. |
![]() | |
|---|---|
| Expert 1 | There is a chest-high obstacle on the building wall. It is a balcony protruding from the wall. Raise your free hand to your head and use it to protect your upper body. Walk forward slowly. |
| Expert 2 | Continue walking directly along the inner guiding line. A trailer hitch protrudes onto the sidewalk from the right. About 5 m ahead, a balcony protrudes slightly to the left. Hold your free hand in front of your face as a protective hand at approximately 8 m. |
| Expert 3 | The sidewalk ahead is narrowed by a person and a trailer hitch protruding far onto the sidewalk. After this narrow section, there is an overhanging balcony at chest height along the inner guide line. Continue straight ahead after that. |
| Gemma4 31B | Continue straight, but veer left to avoid a protruding trailer hitch a few steps ahead at 2 o’clock. Follow the wall on your left for guidance. |
| Gemma4 E2B | Continue straight ahead along the pavement, watch for the curb edge to your right. |
| Gemma4 E2B + LoRa | Continue straight, be careful of the building wall on your right, and watch out for the protruding window frames and the overhang of the balcony above. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Götzelmann, T.; Sommer, S.; Karg, P. Privacy-Preserving Smart Glasses Navigation Through VLM Fine-Tuning for People with Visual Impairments. Electronics 2026, 15, 3223. https://doi.org/10.3390/electronics15143223
Götzelmann T, Sommer S, Karg P. Privacy-Preserving Smart Glasses Navigation Through VLM Fine-Tuning for People with Visual Impairments. Electronics. 2026; 15(14):3223. https://doi.org/10.3390/electronics15143223
Chicago/Turabian StyleGötzelmann, Timo, Sima Sommer, and Pascal Karg. 2026. "Privacy-Preserving Smart Glasses Navigation Through VLM Fine-Tuning for People with Visual Impairments" Electronics 15, no. 14: 3223. https://doi.org/10.3390/electronics15143223
APA StyleGötzelmann, T., Sommer, S., & Karg, P. (2026). Privacy-Preserving Smart Glasses Navigation Through VLM Fine-Tuning for People with Visual Impairments. Electronics, 15(14), 3223. https://doi.org/10.3390/electronics15143223




