⚡️ AI Model Qwen 3.5-Omni Writes Code from Video Guides

Alibaba has released Qwen 3.5-Omni, a new version of the multimodal LLM. The neural network can simultaneously process text, graphic, audio, and video data.

The main difference of Qwen 3.5-Omni is its 256 thousand token context window. This allows the AI to process over 10 hours of audio or approximately 400 seconds of video at 720p resolution at once. Speech recognition covers 113 languages and dialects.
The model was trained on over 100 million hours of audio and video data.

The model "watches" screen recordings with audio instructions and writes working code based on this data without text prompts.

⁉️ This ability emerged accidentally without training)

Get it here