UPM Institutional Repository

Multi-model expert–AI triangulated validation of highway landscape classification with zero-shot multimodal large language models


Citation

Gao, Hangyu and Smardon, Richard and Abu Bakar, Shamsul and Maulan, Suhardi and Yang, Jiani and Mundher, Riyadh and Zhou, Yijiang (2026) Multi-model expert–AI triangulated validation of highway landscape classification with zero-shot multimodal large language models. Journal of Asian Architecture and Building Engineering. pp. 1-25. ISSN 1346-7581; eISSN: 1347-2852 (In Press)

Abstract

Highway landscapes influence driver safety, cultural identity, and economic development, yet existing classification approaches rarely quantify label reliability or validate AI automation against expert judgment. This study develops a multi-model expert–AI triangulated validation framework for highway landscape classification using multimodal large language models (MLLMs). From 1,828 images sampled at 250 m intervals along 418 km of Malaysia’s North–South Expressway, zero-shot classifications from ChatGPT-5.4 Thinking, Claude Sonnet 4.6, and Gemini 3 Thinking were compared with expert annotations across 16 landscape characters. A seven-category agreement scheme yielded 1,352 high-confidence images (74.0% retention). Claude achieved the highest accuracy (70.3%), followed by ChatGPT (68.8%) and Gemini (67.0%); only the Claude–Gemini gap was significant (p < 0.05). All models exceeded 90% retention for morphologically distinct landscapes but disagreed on transitional characters, revealing fragile taxonomic boundaries rather than model error alone. Complementary strengths—Claude in infrastructure-heavy environments, ChatGPT in built-environment characters, Gemini in geomorphologically distinctive features—support triangulation over single-model reliance. The framework reframes validation as mapping a “reliability surface,” treating MLLMs as additional raters rather than replacements, and the validated dataset supports corridor-scale mapping and AI–human collaboration in environmental assessment.


Download File

[img] Text
127672.pdf - Published Version
Available under License Creative Commons Attribution.

Download (9MB)

Additional Metadata

Item Type: Article
Subject: Civil and Structural Engineering
Subject: Architecture
Subject: Cultural Studies
Divisions: Faculty of Design and Architecture
DOI Number: https://doi.org/10.1080/13467581.2026.2706274
Publisher: Taylor and Francis Ltd.
Keywords: AI-expert validation framework; Highway landscape classification; landscape character; multimodal large language models; transportation planning
Sustainable Development Goals (SDGs): SDG 9: Industry, Innovation and Infrastructure
Depositing User: Ms. Siti Radziah Mohamed@mahmod
Date Deposited: 06 Aug 2026 01:30
Last Modified: 06 Aug 2026 01:30
Altmetrics: http://www.altmetric.com/details.php?domain=psasir.upm.edu.my&doi=10.1080/13467581.2026.2706274
URI: http://psasir.upm.edu.my/id/eprint/127672
Statistic Details: View Download Statistic

Actions (login required)

View Item View Item