Evaluation Report · 8 August 2026
A comparative visual study using Xin Qiji’s Water Dragon Chant as a shared literary reference.
This report evaluates how three generative image systems represent a classical ci, 《水龍吟·登建康賞心亭》 (Water Dragon Chant · Ascending the Xīn Pavilion in Jiankang). The evaluation considers text rendering, visual composition, atmosphere, and the ability to revise individual elements while preserving the overall scene.
The results are indicative rather than statistically controlled. In this test, ChatGPT produced the most accurate text rendering, while Grok Imagine Image 2.0 showed strong composition, atmosphere, and editability. Gemini produced a less successful result for this particular prompt.
The reference text is associated with 立秋 (Lìqiū), one of the 24 traditional solar terms. The term means “Beginning of Autumn” and marks the seasonal transition into autumn.
The selected author is Xin Qiji (辛棄疾, 1140–1207), an influential literary figure of the Southern Song period. The selected work is a cí (詞) rather than a shī (詩): a classical lyric form shaped by established tonal and rhythmic patterns.
The same literary reference was used to assess three systems: Grok Imagine Image 2.0, ChatGPT, and Gemini. The outputs were reviewed qualitatively against four criteria:
Text accuracy: the legibility and correctness of the Traditional characters rendered in the image.
Visual interpretation: the relationship between the generated landscape and the mood of the ci.
Composition: the arrangement of mountains, water, architecture, figures, light, and other visual elements.
Editability: the ability to revise individual components without losing the coherence of the original composition.
Approximately 85% of the rendered text appeared correct in the Grok output. This is a strong result for a prompt containing dense Traditional characters and classical literary language.
The system also produced more than a generic mountain background. The scene included distant mountains, autumn water, a pavilion, a scholar, a warrior, a setting sun, and birds. Together, these elements suggested the ci’s combination of ambition, distance, and frustration.
Grok’s most notable strength in this test was the ability to revise individual elements without completely disrupting the composition.
The scene was subsequently revised into a color version while retaining its main spatial relationships.
The revised scene was also converted into a video, extending the experiment from static image generation to animated presentation.
ChatGPT · approximately 95% text accuracy. ChatGPT produced the most accurate Traditional-character rendering in this comparison while also producing a visually coherent result.
Grok Imagine Image 2.0 · approximately 85% text accuracy. Its composition, atmosphere, and ability to keep editing the scene were the strongest aspects of its result.
Gemini. Its output was less successful than the other two systems in this particular test, especially in the combined treatment of text and visual composition.
This comparison is based on a small qualitative sample rather than a controlled benchmark. The accuracy estimates are observational, and results may change with different prompts, model versions, image settings, or evaluation criteria.
Even with these limitations, the experiment shows how a ci from the Mandarin literary tradition, written nearly 900 years ago, can become something that can be viewed, edited, and set in motion. Generative image systems are therefore useful not only for producing illustrations, but also for exploring how literary mood and imagery can be translated into visual form.
辛棄疾 · Xin Qiji · Southern Song Dynasty