Citation
Gao, Hangyu and Smardon, Richard and Abu Bakar, Shamsul and Maulan, Suhardi and Yang, Jiani and Mundher, Riyadh and Zhou, Yijiang
(2026)
Multi-model expert–AI triangulated validation of highway landscape classification with zero-shot multimodal large language models.
Journal of Asian Architecture and Building Engineering.
pp. 1-25.
ISSN 1346-7581; eISSN: 1347-2852
(In Press)
Abstract
Highway landscapes influence driver safety, cultural identity, and economic development, yet existing classification approaches rarely quantify label reliability or validate AI automation against expert judgment. This study develops a multi-model expert–AI triangulated validation framework for highway landscape classification using multimodal large language models (MLLMs). From 1,828 images sampled at 250 m intervals along 418 km of Malaysia’s North–South Expressway, zero-shot classifications from ChatGPT-5.4 Thinking, Claude Sonnet 4.6, and Gemini 3 Thinking were compared with expert annotations across 16 landscape characters. A seven-category agreement scheme yielded 1,352 high-confidence images (74.0% retention). Claude achieved the highest accuracy (70.3%), followed by ChatGPT (68.8%) and Gemini (67.0%); only the Claude–Gemini gap was significant (p < 0.05). All models exceeded 90% retention for morphologically distinct landscapes but disagreed on transitional characters, revealing fragile taxonomic boundaries rather than model error alone. Complementary strengths—Claude in infrastructure-heavy environments, ChatGPT in built-environment characters, Gemini in geomorphologically distinctive features—support triangulation over single-model reliance. The framework reframes validation as mapping a “reliability surface,” treating MLLMs as additional raters rather than replacements, and the validated dataset supports corridor-scale mapping and AI–human collaboration in environmental assessment.
Download File
Additional Metadata
Actions (login required)
 |
View Item |