Narrative design in historical games plays a critical role in shaping player immersion, cultural understanding, and historical representation. With the rapid advancement of large language models, their potential to generate historically grounded interactive narratives has drawn increasing attention, yet systematic comparisons with human-authored content remain scarce. This study investigates the narrative performance of AI models and professional game designers across three culturally distinct historical settings: Ancient Egypt, Medieval Europe, and Ming Dynasty China. A two-stage empirical design was employed: narrative generation by 10 advanced AI models and 30 designers, followed by expert evaluation using a validated five-dimensional framework encompassing historical authenticity, interactivity, immersion, structure, and innovation. The iterative development of the five-dimensional framework makes it a reusable methodological tool. Structural equation modeling revealed that AI excelled in structural coherence and adaptability, especially in culturally distant contexts, while human designers achieved greater cultural depth and emotional resonance in familiar settings. Results indicate that narrative quality does not depend solely on authorship but on the interaction between cultural familiarity, domain expertise, and narrative dimensions. The findings highlight the complementary strengths of AI and human creators, supporting hybrid workflows that combine AI’s efficiency with human cultural insight. Cultural familiarity was identified as a key moderating variable. Statistical analysis indicates that approximately 70% of the evaluated models achieved or exceeded the average performance of human designers in overall scores. However, this improvement was primarily driven by higher performance in the structure and interactivity dimensions. When these technical dimensions were excluded, only about 30% of frontier-scale models exhibited performance comparable to that of human experts in cultural authenticity.