
Today we’re launching LunaRoute.
We’ve been working on it with our early users for the past several months, and I’m very happy to finally open it up. The simplest way to describe it is private AI capacity at a fixed cost. But what I really want to talk about is what you get to do with it.
Room to experiment
I like building things, and quite a few of them start with a fairly vague idea that something might be possible. I don’t necessarily know what I’ll use it for, or whether I’ll keep using it after I’ve built it. Sometimes I just want to understand how something works. Having to figure out in advance whether the result will be useful would take a lot of the fun out of it.
AI has made this a lot more interesting. There are projects I would previously have left alone because getting to a first working version would take more time than I had. Now I can start playing with an idea, get something running, and see whether there’s anything there. And once something works, even badly, it’s much easier to figure out what to try next.
That’s the part I enjoy. You get something working and immediately want to see what else you can do with it. Or it doesn’t work, but you notice something in the result that gives you another idea. You change a few things, run it again, and end up spending time on something you hadn’t even thought about when you started.
When I’m in the middle of that, I want to see where it goes. But when every attempt adds to the bill, I start wondering how much I’ve spent and whether I should keep going. That’s an annoying reason to stop working on something you’re still curious about.
There’s a particular freedom in being able to keep going without that worry. You can try the more complicated approach just to see whether it works better. You can run something across all your data because you’re curious about what you’ll find. You can build a tool that only you will ever use, for a problem that probably didn’t need quite this much attention. I happen to like those projects.
And you can get it wrong. An experiment that didn’t work doesn’t need to be followed by a bill that makes you regret trying it.
When you’re token scarce, you start making decisions around it, sometimes without really noticing. You give the model less information. You try fewer things. You decide that an idea probably isn’t worth running, even though running it is how you’d find out.
I’m interested in what people start building once they’ve stopped worrying about the cost of every attempt. After a while, do you start trying things you wouldn’t even have considered before? That’s what I’m excited to find out.
This matters inside companies too. Someone has an idea for doing something useful with years of documents, or thinks an agent could take care of an annoying part of their job. They might need a few days of trying things before they can explain exactly what it should become. I’d like them to have that room without needing to estimate a token budget for something they haven’t figured out yet.
That’s a big part of what we want to make possible with LunaRoute.
AI capacity at a fixed cost
So we built LunaRoute around capacity. You pay a fixed cost and use that capacity, rather than paying separately for every input and output token. The hardware still has limits, of course. There’s only so much it can do at once. But trying another approach doesn’t add another set of token charges to your bill.
There’s a fair amount of scheduling magic under the hood that lets us get more useful work out of the same GPUs. I’ll write about that separately. You shouldn’t have to understand any of it to use the service.
We run the models and operate the servers ourselves. We’re responsible for keeping them running and figuring out how to make them work better. I enjoyed building my own GPU rig, but I don’t think that should be a prerequisite for having the freedom to experiment with AI.
Your data stays yours
For a company, there’s another thing that can stop an experiment before it starts: whether you’re comfortable giving the service your actual data. You can only learn so much from testing with made-up documents. At some point, you need to use the information that would actually make the experiment useful.
That’s why our hosted inference has Zero Data Retention by default. We don’t retain your prompts or model responses after processing your requests. When you use AI to work through private information, that information shouldn’t become a collection of data sitting with someone else afterward.
The tools around the models
We’re also building out the tools around the models. Once you start putting these things to work, you quickly need more than inference. There’s a scanned document to read, a file to convert, something to look up on the web, or an image to generate and then change a few times until it’s right.
LunaRoute brings image generation and editing, document conversion, AI OCR, and web search together with the models. We’re building an ecosystem you can use to put things together, without having to spend the first part of every experiment connecting another handful of services. You should be able to follow an idea even when it turns out to need more than a text response.
We also have a few Jev-style decision models, including Djev, for classifying requests, choosing an agent’s next step, or scoring a result. And we have a few embedding models for finding related information across your documents.
We’re just getting started
I think AI is one of the most important tools we’ll get to use, and we’re still very early in figuring out what to do with it. A lot of that figuring out is going to happen through people playing with things, trying something that sounds a little unreasonable, and seeing what happens. I want us to make that easier for individual builders and for the people trying to change how their organizations work.
We’ve spent the past several months getting feedback, fixing problems, and tuning the service around the ways people actually use it. Thank you to everyone who has been building with us, telling us when things broke, and coming back to try again. You’ve helped us get it here.
A special thank-you to Jesse Vincent, Harper Reed, Charlie Graham, Braydon McCormick, Dan Shapiro, Dave Morin, Dave Sifry, David Gutelius, Drew Breunig, Erik Garrison, Justin McCarthy, Dan Pupius, Will Smith, and all of our other early users. You’ve pushed the models and the service hard, helping us find the rough edges and make LunaRoute better. If I’ve missed you, ping me and I’ll add you :)
I also want to thank our advisors, Jesse Robbins and Joey Simhon, for their guidance and support as we’ve built LunaRoute.
And thank you to our investors, Offline Ventures, Bloomberg Beta, Octave Fund, Entropy Ventures, and our angel investors, for believing in what we’re building and helping us get here.
LunaRoute is open now. Take that idea you’ve been putting off, give it a proper try, and tell us how it went. I’m looking forward to seeing what you build.