SAGE-IoT: A Lightweight LLM-Inspired Framework for Multimodal Fusion and Bayesian Decision-Making in IoT
DOI:
https://doi.org/10.71146/kjmr1007Keywords:
Edge Intelligence, Multimodal Data Fusion, Parameter-Efficient Fine Tuning, Risk Aware Decision MakingAbstract
This paper proposes SAGE-IoT, a lightweight, large-language-model-inspired multimodal decision-making framework for addressing two persistent challenges in Internet of Things (IoT) scenarios: the difficulty of jointly modeling logs, images, and sensor time series, and the asymmetric costs of decision errors. Drawing on the ideas behind Parameter-Efficient Fine-Tuning (PEFT), Knowledge-Augmented Reasoning (KAR), and hierarchical prompting, the framework implements a “multimodal encoding → gated fusion → knowledge-enhanced reasoning → Bayesian action selection → closed-loop feedback” pipeline that strengthens edge decision-making without deploying a complete Large Language Model (LLM). Experiments across three simulated scenarios, smart home, industrial fault warning, and smart transportation, achieve accuracy of 94.6%, 88.4%, and 91.9%, respectively. Ablation results show that KAR and PEFT contribute gains of approximately 3.8 and 1.7 percentage points, respectively. These results indicate that lightweight LLM-inspired mechanisms offer a feasible way to enhance traditional multimodal models, and that “calibrate before deciding” should be a basic design principle for intelligent IoT decision-making systems.
Downloads
References
[1] A. Giuliano, A. McCafferty-Leroux, J. Yawney, and S. Andrew Gadsden, “Cognitive Internet of Things: A Review of Theory, Applications, and Recent Advances,” IEEE Commun. Surv. Tutor., vol. 28, pp. 446–484, 2026, doi: 10.1109/COMST.2025.3615461.
[2] T. Sadiq and C. W. Omlin, “Sensing in Smart Cities: A Multimodal Machine Learning Perspective,” Smart Cities, vol. 9, no. 1, p. 3, Jan. 2026, doi: 10.3390/smartcities9010003.
[3] K. Taha, “Big Data Analytics in IoT, social media, NLP, and information security: trends, challenges, and applications,” J. Big Data, vol. 12, no. 1, p. 150, Jun. 2025, doi: 10.1186/s40537-025-01192-9.
[4] O. B. Sezer, E. Dogdu, and A. M. Ozbayoglu, “Context-Aware Computing, Learning, and Big Data in Internet of Things: A Survey,” IEEE Internet Things J., vol. 5, no. 1, pp. 1–27, Feb. 2018, doi: 10.1109/JIOT.2017.2773600.
[5] G. Qu, Q. Chen, W. Wei, Z. Lin, X. Chen, and K. Huang, “Mobile Edge Intelligence for Large Language Models: A Contemporary Survey,” IEEE Commun. Surv. Tutor., vol. 27, no. 6, pp. 3820–3860, Dec. 2025, doi: 10.1109/COMST.2025.3527641.
[6] A. Sharshar, L. U. Khan, W. Ullah, and M. Guizani, “Vision-Language Models for Edge Networks: A Comprehensive Survey,” IEEE Internet Things J., vol. 12, no. 16, pp. 32701–32724, Aug. 2025, doi: 10.1109/JIOT.2025.3579032.
[7] J.-S. Yang, Z. Shen, Z. Zeng, and Z. Chen, “Domain-Adapted Large Language Models for Industrial Applications: From Fine-Tuning to Real-Time Deployment,” Comput. Sci. Bull., vol. 8, no. 01, pp. 272–289, Dec. 2025, doi: 10.71465/csb162.
[8] F. Sarhaddi et al., “LLMs and IoT: A Comprehensive Survey on Large Language Models and the Internet of Things,” IEEE Open J. Commun. Soc., vol. 7, pp. 3585–3616, 2026, doi: 10.1109/OJCOMS.2026.3680064.
[9] H. Xu, L. Han, Q. Yang, and M. Li, “Penetrative AI: Making LLMs Comprehend the Physical World | Proceedings of the 25th International Workshop on Mobile Computing Systems and Applications,” ACM Conferences. Accessed: Sep. 10, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3638550.3641130
[10] A. Liu, W. Jiang, and Z. Feng“Multi-Modal Integrated Sensing and Communication in Internet of Things With Large Language Models.” Accessed: Sep. 10, 2026. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/11026879
[11] T. An, Y. Zhou, H. Zou, and J. Yang, “IoT-LLM: A framework for enhancing large language model reasoning from real-world sensor data,” Patterns, vol. 7, no. 1, Jan. 2026, doi: 10.1016/j.patter.2025.101429.
[12] D. Jiang, Z. Shen, Q. Zheng and J. Jin“Farm-LightSeek: An Edge-Centric Multimodal Agricultural IoT Data Analytics Framework With Lightweight LLMs.” Accessed: Sep. 10, 2026. [Online]. Available: https://ieeexplore.ieee.org/abstract/document/11027609
[13] M. K, R. G, A. Mondal, S. Singh, S. Tiwari, and J. Sharma, “HARMONY: A Framework for Multimodal LLM-Powered AI Agents in Smart Homes via the Model Context Protocol,” IEEE Access, vol. 14, pp. 13669–13687, 2026, doi: 10.1109/ACCESS.2026.3653992.
[14] H. Zhang, M. Farzanullah, M. Ghassemi, A. Bin Sediq, A. Afana, and M. Erol-Kantarci, “Multi-Modal Data-Enhanced Foundation Models for Prediction and Control in Wireless Networks: A Survey,” IEEE Commun. Surv. Tutor., vol. 28, pp. 4359–4393, 2026, doi: 10.1109/COMST.2025.3648785.
[15] J. Li, J. Li, G.. Yang, H. Chi and L. Yang“Applications of Large Language Models and Multimodal Large Models in Autonomous Driving: A Comprehensive Review.” Accessed: Sep. 10, 2026. [Online]. Available: https://www.mdpi.com/2504-446X/9/4/238
[16] C. Zhang, Z. Yang, X. He, and L. Deng, “Multimodal Intelligence: Representation Learning, Information Fusion, and Applications,” IEEE J. Sel. Top. Signal Process., vol. 14, no. 3, pp. 478–493, Mar. 2020, doi: 10.1109/JSTSP.2020.2987728.
[17] W. Shao, D. Fan, C. Cui, Y. Xu, and X. Lyu“Deep Multimodal Data Fusion,” ACM Comput. Surv., Accessed: Sep. 10, 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3649447
[18] P. Wang, S. Song, H. Ji, S. Cao, and J. Hu“From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning | OpenReview.” Accessed: Sep. 10, 2026. [Online]. Available: https://openreview.net/forum?id=yfTU8FTS2Z
[19] S. Luo et al., “Delving Into Multi-Modal Multi-Task Foundation Models for Road Scene Understanding: From Learning Paradigm Perspectives,” IEEE Trans. Intell. Veh., vol. 9, no. 12, pp. 8040–8063, Dec. 2024, doi: 10.1109/TIV.2024.3406372.
[20] T. Jiao, C. Guo, X. Feng, Y. Chen, and J. Song, “A Comprehensive Survey on Deep Learning Multi-Modal Fusion: Methods, Technologies and Applications,” Comput. Mater. Contin., vol. 80, no. 1, pp. 1–35, Jul. 2024, doi: 10.32604/cmc.2024.053204.
[21] S. Lim, H.-C. Kwon, and M. Kim, “When to Invoke an LLM in Industrial Natural-Language-to-DSL Conversion: A Taxonomy-Driven, Confidence-Gated Selective Hybrid for Safety-Critical Avionics Test-Script Authoring,” IEEE Access, pp. 1–1, 2026, doi: 10.1109/ACCESS.2026.3730474.
[22] C. Xu et al., “LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection,” in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens, Eds., San Diego, California, United States: Association for Computational Linguistics, Jul. 2026, pp. 11881–11910. doi: 10.18653/v1/2026.acl-long.546.
[23] C. Xu et al., “LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection,” in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens, Eds., San Diego, California, United States: Association for Computational Linguistics, Jul. 2026, pp. 11881–11910. doi: 10.18653/v1/2026.acl-long.546.
[24] F. Yang, X. Zhang and Z. Yu“Human–Machine Collaborative Learning for Streaming Data-Driven Scenarios.” Accessed: Sep. 10, 2026. [Online]. Available: https://www.mdpi.com/1424-8220/25/21/6505
Downloads
Published
License
Copyright (c) 2026 Qifan Shen, Qirong Shen, Lei Sun, Sihan Shen, Yutong Ye, Muhammad Wahab Hanif (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.
