MiniCPM-V 4.5 β Overview
A vision-language MLLM with strong performance on image / video understanding, OCR, and document parsing. Released alongside MiniCPM-o 4.5.
Looking at this for a new project? Consider MiniCPM-V 4.6 instead β it has a longer context, faster vision tower, and is the actively-recommended release.
Highlights
- 9B parameters, Qwen3 backbone
- Image, multi-image, and video understanding (no audio)
- One model with optional thinking mode (toggle via
enable_thinking) - HuggingFace:
openbmb/MiniCPM-V-4_5Β· ModelScope:OpenBMB/MiniCPM-V-4_5