Thinking Machines drops Inkling, a 975B MoE open-weights model with a surprising pitch: 'We aren't the strongest, but we are the most adaptable.' What's the catch?

Another trillion-parameter beast enters the AI colosseum, but instead of the usual "we absolutely crushed GPT-4" benchmark flexing, the creators openly admit: "Yeah, we aren't the strongest model out there right now." Bold strategy. Let's see if it pays off.
Created by Thinking Machines, Inkling is making waves in the dev community. Grab your coffee, sit down, and let's dissect whether this open-weights model is a genuine gift to open-source or just a high-key marketing funnel for their paid platform.
Thinking Machines just released the weights for Inkling, their first open-weights model under the Apache 2.0 license. Hardware-wise, it's a massive heavy-lifter:
But the real kicker isn't the specs; it's the philosophy. They are incredibly upfront that Inkling isn't designed to sit at the top of temporary leaderboards. Instead, they pitched it as a highly moldable foundation. To prove this, they literally made Inkling write and execute its own fine-tuning job to transform itself into a model that completely avoids using the letter "e".
As soon as the model hit Product Hunt, the community split into very distinct camps.
Most devs loved the humble approach. In an industry bloated with overhyped benchmarks, hearing a lab say "we aren't the strongest, but we are the most adaptable" is incredibly refreshing. Getting a 1M context multimodal MoE under Apache 2.0 is a massive win for open-source advocates.
On the other side, veteran systems architects quickly pointed out the elephant in the server room: 975B is absolutely massive.
One developer noted: "Let's be real. 'Open weights' and 'actually touchable' are two different things at 975B. Who is actually running this locally? Most people don't have supercomputer clusters in their basement. In reality, the open-source aspect is just a trust signal to lure you into their managed platform, Tinker, which will be the only realistic way most devs can interact with it."
Unless you have deep pockets to rent massive GPU clusters, you won't be running this on a basic setup.
Those who actually got their hands dirty had mixed experiences. Some tested a quick LoRA job on Tinker and were pleasantly surprised that it worked smoothly right out of the box without the usual GPU setup dance.
However, power users are demanding more:
At the end of the day, Thinking Machines' approach is pragmatism at its finest. The AI hype train is slowly shifting away from giant generalist models to custom, domain-specific tooling.
The takeaway for devs:
Reference: Product Hunt