MiniCPM-o 4.5 β Overview
The omni-modal flagship of the MiniCPM-o series β vision, speech, and full-duplex live streaming, all in a single model.
Highlights
- End-to-end omnimodal β vision encoder + speech encoder + LLM + TTS densely connected via hidden states
- Full-duplex live streaming β hear, see, and speak simultaneously
- Voice cloning and configurable assistant voice
- HuggingFace:
openbmb/MiniCPM-o-4_5Β· ModelScope:OpenBMB/MiniCPM-o-4_5
For real-time, full-duplex demos see the MiniCPM-o-Demo repository β it ships a turn-key web demo for the o-series with PyTorch + CUDA inference and a frontend.
Where to start
- Audio recipes: Speech-to-Text, Text-to-Speech, Voice Cloning
- Deploy: vLLM, SGLang, llama.cpp, Ollama
- Quantize: GGUF