Inkling: Thinking Machines' 975B Open Multimodal Model Signals a New Wave in Video+Audio AI | Cybernomics
researchWednesday, July 15, 2026

Inkling: Thinking Machines' 975B Open Multimodal Model Signals a New Wave in Video+Audio AI

Thinking Machines Lab released Inkling, a 975-billion-parameter open-source model trained to understand video and audio, positioning the lab as a serious competitor in multimodal AI. Inkling's openness and media focus accelerate capabilities for video understanding, search, and automation while raising new operational, safety, and infrastructure questions for enterprises.

Inkling represents a notable moment in model development: a very large, multimodal model released openly with explicit grounding in video and audio modalities. Unlike many text-first models, Inkling is trained to reason about temporal, visual, and auditory signals-making it directly applicable to media indexing, automated captioning, action recognition, and multimodal search. Open-source availability lowers barriers for startups and research groups to experiment, fine-tune, and build differentiated products without the guardrails of closed ecosystems.

For businesses, the practical implications are immediate. Media-heavy industries-entertainment, surveillance, retail, and manufacturing-can prototype advanced workflows such as automated content summarization, searchable video archives, and real-time audio analytics. However, adopting a 975B model introduces heavy compute demands and integration complexity: fine-tuning, serving, and embedding such a model into production pipelines requires substantial GPU/TPU capacity, MLOps maturity, and cost modeling.

Inkling's openness also brings governance and safety considerations. Multimodal models can expose sensitive information (faces, voices, proprietary content), so leaders should implement strong data handling policies, model evaluation for bias and hallucination in visual/audio contexts, and red-team testing specific to media misuse. Additionally, licensing and downstream IP questions arise when models are trained on third-party video/audio.

Actionably, businesses should run targeted pilots focused on high-value use cases where video/audio understanding drives measurable ROI, invest in infrastructure or cloud partnerships that support large-scale multimodal models, and engage legal and security teams early. Firms that pair experimentation with disciplined governance and cost controls will capture first-mover advantages in media-centric AI automation.

multimodalopen-sourcevideo-aimodels

Original Source

WIRED

Read Original