The open-weight model supports text, image and audio reasoning with a context window of up to one million tokens.