What it is
The job it was built for, and the job it is not for.
GLM 4.6V is a large multimodal model whose job is visual understanding and long-context reasoning over mixed media: images, documents and video together with text. It is built to handle complex page layouts rather than to act as a plain chat assistant, so the natural work is extraction, description and reasoning about what a page or a frame actually contains. It is not a text-only writing model, and there is nothing in Z.ai's description to suggest it generates images or video of its own.