Please allow me to quote a recent response regarding the billing structure:
Beyond the billing mechanics mentioned above, I want to share a few more realities regarding the system itself that are contributing to this situation:
-
The Reality of Model Scaling: The K3 model operates at a massive 2.8T scale. Moving from a 1T to a 2.8T model does not simply mean a 2.8x increase in required inference resources—the resource demand scales exponentially.
-
Hardware Bottlenecks: At the same time, expanding our GPU capacity is exceptionally difficult in the current hardware climate. As a result, our service is currently operating in a highly saturated state.
-
What We Are Doing: We are attacking this problem from two angles. First, we are pushing as hard as possible to expand our physical capacity. Second, our inference service developers are deeply focused on improving the model’s underlying inference efficiency. We already have clear, actionable directions to boost this efficiency further in the near future.
-
Rethinking Quota Assumptions: We are learning that our users interact with our models in incredibly diverse ways. Our previous quota design was built on the assumption that “everyone uses a little bit of everything everywhere.” This assumption has proven flawed, creating a contradictory design that restricts users like you. Fixing this is a primary focus for us right now.
Thank you again for believing in Kimi enough to invest your time and money into it. We are working hard to build a plan structure that respects that investment. We just ask that you give us a little more time to get these improvements implemented.
With sincere gratitude,
Yu