Splat Search: Approach toward Real-Time Semantic Digital Twin
Keywords:
Digital twins, Gaussian splatting, Natural language search, Real-Time 3d reconstruction, RGB-D mapping, Semantic SLAMAbstract
To meet the demand for high-fidelity spatial awareness in robotics and Augmented Reality (AR), this project introduces a unified framework for generating real-time, linguistically-aware digital twins. Traditional Simultaneous Localization and Mapping (SLAM) systems typically suffer from "semantic blindness"—they capture 3D geometry but lack object context. This system bridges the gap between spatial data and human language by integrating SplaTAM for dense Gaussian Splatting, LangSplat for open-vocabulary embedding, and Online LangSplat for high-speed encoding. The architecture leverages a robust client-server model where an iPhone streams synchronized RGB-D data directly to a GPU-accelerated backend. This setup enables real-time 3D reconstruction at over 45 frames per second. During this generation phase, the system dynamically assigns CLIP-based language features to millions of 3D Gaussians. The resulting digital twin is not only photorealistic but fully searchable via natural language queries. Users can seamlessly localize and identify specific objects within previously unknown environments with remarkable semantic accuracy. By eliminating the latency gap between mapping and understanding, this research delivers a scalable, highly efficient solution perfectly suited for advanced smart-home automation and assistive navigation technologies.
References
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3D Gaussian splatting for real-time radiance field rendering,” ACM Transactions on Graphics, vol. 42, no. 4, pp. 1–14, Jul. 2023.
A. Pumarola, E. Corona, G. Pons-Moll, and F. Moreno-Noguer, “D-NeRF: Neural radiance fields for dynamic scenes,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 2021, pp. 10313-10322.
T. Müller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,” ACM Transactions on Graphics, vol. 41, no. 4, pp. 1–15, Jul. 2022.
G. Wu et al., “4D Gaussian splatting for real-time dynamic scene rendering,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 20310–20320.
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik, “LERF: Language embedded radiance fields,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 19672-19682.
A. Radford et al., “Learning transferable visual models from natural language supervision,” In International Conference on Machine Learning (ICML), 2021, pp. 8748–8763.
K. Liu, F. Zhan, J. Zhang, M. Xu, Y. Yu, A. EI Saddik, C. Theobalt, E. Xing, and S. Lu, “Weakly supervised 3D open-vocabulary segmentation,” Advances in Neural Information Processing Systems (NeurIPS), 2023, pp. 53433-53456.
J.-C. Shi, M. Wang, H.-B. Duan, and S.-H. Guan, “Language embedded 3D Gaussians for open-vocabulary scene understanding,” In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 5333–5343.
M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister, “LangSplat: 3D language Gaussian splatting,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 20051-20060.
A. Kirillov et al., “Segment anything,” 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 3992-4003.
J. Cen et al., “Segment anything in 3D with NeRFs,” Proceedings of the 37th International Conference on Neural Information Processing Systems, 2023, pp. 25971–25990.
C. M. Kim, M. Wu, J. Kerr, K. Goldberg, M. Tancik, and A. Kanazawa, “GARField: Group anything with radiance fields,” 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 2024, pp. 21530-21539.
Y. Ji, Y. Liu, G. Xie, B. Ma, Z. Xie, and H. Liu, “NEDS-SLAM: A neural explicit dense semantic SLAM framework using 3D Gaussian splatting,” in IEEE Robotics and Automation Letters, vol. 9, no. 10, pp. 8778-8785, Oct. 2024.
Y. Liu, L. Kong, J. Cen, R. Chen, W. Zhang, L. Pan, K. Chen, and Z. Liu, “Segment any point cloud sequences by distilling vision foundation models,” arXiv preprint arXiv:2306.09347, pp. 37193-37229, 2023.
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled feature fields enable few-shot language-guided manipulation,” arXiv preprint arXiv:2308.07931, 2023.
Y. Siddiqui, L. Porzi, S. Rota Bulo, N. Müller, M. Nießner, A. Dai, and P. Kontschieder, “Panoptic lifting for 3D scene understanding with neural fields,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 2023, pp. 9043-9052.
S. Zhi, T. Laidlow, S. Leutenegger, and A. J. Davison, “In-place scene labelling and understanding with implicit scene representation,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 2021, pp. 15838–15847.