Globe Lens: An AI-powered Multimodal Travel Assistance System for International Travellers

Authors

  • Karanam Rohitha Amrutha Varshini
  • Kamisetti Pragna Sri Sai Lakshmi
  • Shaik Haneefa
  • Darapu Uma

Keywords:

Computer vision, Intelligent tourism systems, Large language models, Multimodal AI, Natural language processing, Real-time translation, Travel assistance system

Abstract

The language barrier, cultural diversity, real-time navigation requirements, and the general plethora of local customs and regulations among countries have complicated international travel. Conventional traveling support devices tend to work in isolation—providing translation, maps or suggestions, but seldom combine these capabilities into a single, smart experience. Globe Lens is a proposed AI-based multimodal travel assistance system that will be a robust global traveller digital assistant. The system uses large language models (LLMs), computer vision, speech recognition, and real-time geospatial information to offer context-sensitive, on-command assistance in various modalities such as text, voice and image inputs. Its core functions are real-time visual translation of signs and menus using the camera, voice-activated conversation in 100-plus languages, AI-powered itinerary planning based upon the tastes and preferences of the user and the real-time environment, cultural etiquette advice, emergency response navigation, and currency and unit conversion. Globe Lens will use a modular microservices model implemented on cloud infrastructure that allows scaling and provides offline backup services in low-connectivity areas. Early system tests reveal that there is high accuracy in object-based translating exercises and high ratings of user satisfaction in simulated travel situations. The study includes the system design, methodology, literature context, and discussion of results and future research directions.

References

R. Burke, “Hybrid recommender systems: Survey and experiments,” User Modeling and User-Adapted Interaction, vol. 12, no. 4, pp. 331–370, Nov. 2002.

Y. Wang, S. T. Xia, and J. Wu, “A less-greedy two-term Tsallis entropy information metric approach for decision tree classification,” Knowledge-Based Systems, vol. 120, pp. 34–42, Mar. 2017.

J. Borràs, A. Moreno, and A. Valls, “Intelligent tourism recommender systems: A survey,” Expert Systems with Applications, vol. 41, no. 16, pp. 7370–7389, Nov. 2014.

J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 2019, pp. 4171–4186.

D. Koehn, S. Lessmann, and M. Schaal, “Predicting online shopping behaviour from clickstream data using deep learning,” Expert Systems with Applications, vol. 150, p. 113342, Jul. 2020.

T. Brown et al., “Language models are few-shot learners,” Advances in Neural Information Processing Systems, vol. 33, pp. 1877–1901, 2020.

E. Adamopoulou and L. Moussiades, “An overview of chatbot technology,” in Proceedings of the IFIP International Conference on Artificial Intelligence Applications and Innovations, vol. 584, May 2020, pp. 373–383.

B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 11, pp. 2298–2304, Nov. 2017.

X. Zhou, C. Yao, H. Wen, Y. Wang, S. Zhou, W. He, and J. Liang, “EAST: An efficient and accurate scene text detector,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5551–5560.

Shi B, Wang X, Lyu P, Yao C, and Bai X, “Robust Scene Text Recognition with Automatic Rectification,” arXiv preprint arXiv:1603.03915. 2016 Mar 12.

L. Gavalakis and I. Kontoyiannis, “Fundamental limits of lossless data compression with side information,” IEEE Transactions on Information Theory, vol. 67, no. 5, pp. 2680–2692, May 2021.

A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, Oct. 2020.

T. H. Nguyen and K. Shirai, “Topic modelling-based sentiment analysis on social media for stock market prediction,” in Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, Beijing, China, Jul. 2015, pp. 1354–1364.

J. Achiam et al., “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023.

Published

2026-05-30