Porcelain Publishing / IVC / Volume 2 / Issue 3 / DOI: 10.47297/ppiivc2026020306
ARTICLE

Flux in the Folds: The Cross-Media Logic and Embodied Turn of Generative AI Imagery

Tian Zhao1 ,  Xiaoyu Wu1 ,  Yansong Chen1*
Show Less
1 School of Arts and Communication, Beijing Normal University, Beijing 100875, China
Published: 29 September 2026
© 2026 by the Author(s). This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/)
Abstract

With the rapid emergence and in‑depth application of generative artificial intelligence technology, the image medium has gradually shifted from the replication of physical reality under photographic materialism to the probabilistic generation in algorithm‑driven latent spaces. AI imagery completely breaks the indexical connection between traditional images and physical reality, and reshapes the production logic and aesthetic core of images. It not only manifests as the deconstruction and infinite recombination of basic data units such as pixels and feature vectors at the level of material existence, getting rid of the dependence on live‑action equipment, physical scenes and human performers, but also constructs a new logic of sensation free from the constraints of traditional narrative logic at the level of aesthetic expression, allowing the emotional transmission of images to directly act on the audience's senses and body. This paper places AI imagery under the perspective of Spinoza and Deleuze's affect theory. Based on the theoretical framework of affective folds, it takes the cross‑media evolution paths of text‑image‑video, text‑to‑video and text‑to‑multimodal video as clues, and deeply analyzes the affective generation mechanism of AI imagery from three dimensions: cross‑media logic, sensory transmission and embodied viewing. The study argues that the evolution of generative AI imagery enables the deep integration of cross‑media modalities, algorithmic sensation and embodied audiences, thereby forming a new type of "affective machine" imagery centered on affective transmission, which provides a brand‑new thinking path for the development of post‑cinematic image forms in the digital age.

Keywords
AI Imagery
Cross‑Media
Affect Theory
Embodied Viewing
Logic of Sensation
Funding
The Major Program of the National Social Science Fund of China in Arts Studies (Grant No. 25ZD07).
References

[1] Akmeşe, Z. (2025). The aesthetics of imperfection: A theoretical evaluation of the meaning layers of glitch aesthetics in communication design and cinema. CINEJ Cinema Journal, 13(2), 812-844. https://doi.org/10.5195/cinej.2025.905

[2] Balsom, E. (2009). Screening rooms: The movie theatre in/and the gallery. Public: Art/Culture/Ideas, 40, 24-39. https://public.journals.yorku.ca/index.php/public/article/view/31969

[3] Bar-Tal, O., Chefer, H., Tov, O., Herrmann, C., Paiss, R., Zada, S., Ephrat, A., Hur, J., Liu, G., Raj, A., Li, Y., Rubinstein, M., Michaeli, T., Wang, O., Sun, D., Dekel, T., & Mosseri, I. (2024). Lumiere: A space-time diffusion model for video generation. arXiv. https://doi.org/10.48550/arXiv.2401.12945

[4] Bazin, A. (1960). The ontology of the photographic image (Trans. H. Gray). Film Quarterly, 13(4), 4-9. https://doi.org/10.2307/1210183

[5] Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., & Rombach, R. (2023). Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv. https://doi.org/10.48550/arXiv.2311.15127

[6] ByteDance Seed Team. (2026). Seedance 2.0 official launch. https://seed.bytedance.com/en/blog/seedance-2-0-official-launch

[7] Kane, C. L. (2014). Compression aesthetics: Glitch from the avant-garde to Kanye West. InVisible Culture, (21). https://www.proquest.com/docview/1771515644/fulltext?sourcetype=Scholarly%20Journals

[8] Dalsgaard, P. (2025). Thinking through prompting: Cognitive mediation in human–AI interaction. Proceedings of the 36th Annual Conference of the European Association of Cognitive Ergonomics (ECCE '25), 1-6. https://doi.org/10.1145/3746175.3747192

[9] Deleuze, G. (1993). The fold: Leibniz and the baroque (Conley, T., Trans.). University of Minnesota Press.

[10] Deleuze, G. (2003). Francis Bacon: The logic of sensation (Smith, D. W., Trans.). Continuum.

[11] Denson, S. (2023). From sublime awe to abject cringe: On the embodied processing of AI art. Journal of Visual Culture, 22(2), 146-175. https://doi.org/10.1177/14704129231194136

[12] Doane, M. A. (2007). Indexicality: Trace and sign: Introduction. Differences: A Journal of Feminist Cultural Studies, 18(1), 1-6. https://doi.org/10.1215/10407391-2006-020

[13] Gao, H., Qu, J., Tang, J., Bi, B., Liu, Y., Chen, H., Liang, L., Su, L., & Huang, Q. (2025). Exploring hallucination of large multimodal models in video understanding: Benchmark, analysis and mitigation [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.19622

[14] Kling Team. (2025). Kling-Omni technical report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.16776

[15] Kopelman, S., & Frosh, P. (2023). The "algorithmic as if": Computational resurrection and the animation of the dead in Deep Nostalgia. New Media & Society, 27(4), 2393-2413. https://doi.org/10.1177/14614448231210268

[16] Lan, J. (2026). The affective machine: The "new sensibility logic" of generative AI imagery. Film Art, (1), 3-11.

[17] Liu, Z. (2025). Human-AI co-creation: A framework for collaborative design in intelligent systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.17774

[18] Manovich, L. (1995). What is digital cinema? https://manovich.net/index.php/projects/what-is-digital-cinema

[19] Runway Research. (2023). Gen-2: Generate novel videos with text, images or video clips. https://runway.com/research/gen-2

[20] Sahoo, P., Meharia, P., Ghosh, A., Saha, S., Jain, V., & Chadha, A. (2024). A comprehensive survey of hallucination in large language, image, video and audio foundation models. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 11709-11724). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.685

[21] Scott-Stevenson, J., & Wolozin, S. (2024). Embodiment, AI and the perception of the real [Report]. MIT Open Documentary Lab. https://opendoclab.mit.edu/presents/embodiment-ai-perception-real-idfa-mit-research-development-network-report-2024

[22] Spinoza, B. de. (1997). Ethics—Part 3 (Elwes, R. H. M., Trans.). Project Gutenberg. https://www.gutenberg.org/cache/epub/948/pg948-images.html

[23] Wang, H., & Wang, Z. (2025). Research on methods to improve video generation quality based on Google DeepMind's Veo model. In Proceedings of the 2025 5th International Conference on Computer Vision, Application and Algorithm (CVAA) (pp. 274-278). IEEE. https://doi.org/10.1109/CVAA66438.2025.11193318

[24] Xing, Z., Feng, Q., Chen, H., Dai, Q., Hu, H., Xu, H., Wu, Z., & Jiang, Y.-G. (2024). A survey on video diffusion models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2310.10647

[25] Zhou, J., Gao, H., Voleti, V., Vasishta, A., Yao, C.-H., Boss, M., Torr, P., Rupprecht, C., & Jampani, V. (2025). Stable virtual camera: Generative view synthesis with diffusion models [Preprint]. arXiv. https://doi.org/10.48550/arXiv. 2503.14489

[26] Zhou, S., Yang, P., Wang, J., Luo, Y., & Loy, C. C. (2023). Upscale-A-video: Temporal-consistent diffusion model for real-world video super-resolution [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2312.06640

Share
Back to top
Intelligent Visuals and Communication, Electronic ISSN: 2978-5499 Print ISSN: 2978-5480, Published by Porcelain Publishing