Smooth and Flexible Camera Movement Synthesis via Temporal Masked Generative Modeling

Chenghao Xu (Xidian University) · guangtao lyu · Jiexi Yan (Xidian University) · Muli Yang (Xidian University) · Cheng Deng (Xidian University)
camera movement synthesisconsecutive memory encoderdance performancediscrete camera tokenizerdiscrete quantization schemeexperimental evaluationshistorical context modelinglong-term temporal dependenciesmasked token predictiononline camera trajectoriesperformance superiorityreal-time applicationsshort-term temporal dependenciestemporal conditional masked transformertemporal masked generative modeling

In dance performances, choreographers define the visual expression of movement, while cinematographers shape its final presentation through camera work. Consequently, the synthesis of camera movements informed by both music and dance has garnered increasing research interest. While recent advancements have led to notable progress in this area, existing methods predominantly operate in an offline manner—that is, they require access to the entire dance sequence before generating corresponding camera motions. This constraint renders them impractical for real-time applications, particularly in live stage performances, where immediate responsiveness is essential. To address this limitation, we introduce a more practical yet challenging task: online camera movement synthesis, in which camera trajectories must be generated using only the current and preceding segments of dance and music. In this paper, we propose TemMEGA (Temporal Masked Generative Modeling), a unified framework capable of handling both online and offline camera movement generation. TemMEGA consists of three key components. First, a discrete camera tokenizer encodes camera motions as discrete tokens via a discrete quantization scheme. Second, a consecutive memory encoder captures historical context by jointly modeling long- and short-term temporal dependencies across dance and music sequences. Finally, a temporal conditional masked transformer is employed to predict future camera motions by leveraging masked token prediction. Extensive experimental evaluations demonstrate the effectiveness of our TemMEGA, highlighting its superiority in both online and offline camera movement synthesis.