We present GLM-4.5V (106B-A12B) and GLM-4.1V-Thinking (9B), a series of open-sourced VLMs designed to advance general-purpose multimodal understanding and reasoning. With enhanced pre-trained base and carefully optimized multi-domain RL procedure, GLM-4.5V achieves state-of-the-art performance on nearly all tasks among open-source models of similar size in a comprehensive evaluation across 42 public benchmarks.
We introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset contains 6 major capability categories and 21 subcategories, with the video average length of 1.14 hours, approximately four times longer than the longest existing dataset.