Inkling: Open-Weights Multimodal AI for Customizable Intelligence
Inkling is a 975B parameter Mixture-of-Experts transformer with open weights, designed to extend human will and judgment through AI that anyone can customize. Built by Thinking Machines, it offers native multimodal reasoning across text, images, and audio with a 1M token context window, making it a practical foundation for developers and organizations seeking transparent, adaptable AI systems.
Product Highlights
- Open-Weights Architecture: Full model weights available for unrestricted customization, fine-tuning, and deployment without vendor lock-in or hidden constraints.
- Multimodal Native Reasoning: Trained from scratch on 45 trillion tokens of text, images, audio, and video, enabling seamless understanding across modalities without separate encoders.
- Controllable Thinking Effort: Adjustable reasoning intensity lets developers balance performance with cost and latency—achieving comparable results with one-third the tokens of competing models.
- Massive Context Window: Supports up to 1 million tokens for processing extensive documents, long conversations, and complex multi-step workflows.
- Efficient MoE Design: 41B active parameters from 975B total parameters deliver high performance with optimized inference costs.
- Calibrated & Trustworthy: Trained for epistemic reliability—proper confidence calibration, instruction following, and resistance to censorship—enabling dependable real-world deployment.
- Safety-First Training: Strong built-in safeguards validated by external testers, with top-tier refusal rates on harmful requests without over-refusing benign queries.
Use Cases
- Enterprise AI Customization: Fine-tune Inkling on proprietary data through the Tinker platform to create specialized models for finance, legal, healthcare, and other domain-specific applications.
- Agentic Coding & Software Development: Deploy as a base model for autonomous coding agents, with strong performance on SWEBench and Terminal Bench for real-world software engineering tasks.
- Multimodal Content Analysis: Process and reason over combined text, visual, and audio inputs for media analysis, accessibility tools, and interactive applications.
- Real-Time Collaboration Systems: Power voice and vision-enabled interaction models for natural, low-latency human-AI collaboration.
- Forecasting & Prediction: Leverage calibrated uncertainty for prediction markets, risk assessment, and strategic planning where reliable confidence estimation matters.
- Edge & Cost-Sensitive Deployment: Use Inkling-Small (12B active parameters) for latency-critical applications while maintaining competitive performance on reasoning and coding tasks.
Target Audience
Inkling is designed for AI developers, ML engineers, and technical teams at organizations that prioritize model transparency, customization, and multimodal capabilities. It serves those who need to fine-tune foundation models on proprietary data, deploy AI with verifiable safety properties, and maintain control over their AI infrastructure without sacrificing performance.