What it is
The job it was built for, and the job it is not for.
Gemma 3 12B is an open-weight vision-language model: it reads text and images and writes text back. Google positions it for general chat, reasoning and maths work across a wide spread of languages. It does not generate images or audio, and it is not the frontier tier of Google's line-up, so the heaviest research and long-form agentic work belongs elsewhere.