Descripció del projecte
This research proposes a new navigation assistance system for people with low vision, which integrates large language models (LLMs) and advanced localization techniques to improve both scene understanding and spatial accuracy.
The first component focuses on using LLM in conjunction with computer vision to detect, describe, and semanticize visual scenes, providing users with contextually rich and human-like interpretations of their environment in real time. This allows the system to move beyond object detection to meaningful scene understanding, such as identifying navigational possibilities, obstacles, and points of interest.
The second component improves location accuracy by fusing cartographic data, sensor inputs, and multimodal information (e.g., GPS, inertial, and visual cues) to provide precise positioning even in complex urban or indoor environments. Together, these two advances aim to create an intelligent, context-aware navigation system that provides safer, more intuitive, and more personalized mobility support for people with low vision.