NewsLab
Aug 28 12:57 UTC

Show HN: We built open OpenRouter that turns usage into a better model (github.com)

195 points|by SilenN||37 comments|Read full story on github.com
Hi HN, we built an open source model gateway. It's a single place to manage our own self hosted, frontier, and open source models in one place.

It’s is rust native, built for concurrency, and implements all the config quirks across models and providers (streaming formats, tool calls, model parameters, rate limits, and different error behavior).

The gateway adds under 1 ms for BYOK requests and under 2 ms when Experiential supplies the provider key. It has every major inference provider, and 1000+ models refreshed daily via a codex agent that opens a PR.

Compared to other similar projects we’re open source, take no markup, allow you to mix local models with a marketplace, and use your traffic to (opt in) train you a model. Simple routing doesn’t warrant a 10% token markup.

The way we do this is given standardized OTel traces, we mine representative real tasks, use text world models to simulate rollouts for various models, apply an LLM judge, and fit a nearest neighbor classifier on top of an embedding of a prompt to decide the optimal model for each request. Usually this can map out a better pareto curve on cost/quality than just calling single models but it’s not perfect.

Using these simulations we can also do things like suggesting cache hit optimizations, new model suggestions, and training models.

It’s open source, so you can deploy it on your own infrastructure, use our hosted version with 0 markup, or read how we design for maximum availability on our website.

Comments (37)

37 shown
  1. 1. Areibman||context
    Could you say more about how caching works? One major advantage of sticking with a single model is saving money on cached input tokens. I'd imagine if you swap between a bunch of models, you may improve performance but cost would would balloon out of control
  2. 2. purplecats||context
    and caching is related to performance too ofc
  3. 3. SilenN||context
    The trick is to rarely switch, or switch at task boundaries. Often the conclusion of routing is actually "this one model is actually at the pareto front for this task, just use it always".
  4. 4. cameronh90||context
    But then it's better to just not have a gateway switch models at all.

    Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.

  5. 5. SilenN||context
    That is another way to do. Or we can automatically figure out which models the subagents should be using for you. And update them as new models come out and the work your subagents do changes. More than one way to skin a cat.
  6. 6. aerzen||context
    This does make sense. I generally only switch between models in pi when creating a new session. And it is apparent from the promt if this just a "how to see open ports on linux" or "make a concrete plan for feature X"
  7. 7. try-working||context
    Generally you should only have two models in the pool per domain. I wrote some of my learnings building a router here: https://try.works/first-principles-of-model-routing
  8. 8. ashermania||context
    Finally an open source tool doing this!
  9. 9. 23david||context
    Super interesting and congrats on the release. Curious if you initially had this in Python and then rewrote in Rust?
  10. 10. SilenN||context
    Yep! If you look at the commit history that's exactly what happened.
  11. 11. cheema33||context
    I have not tried it yet. Is it similar to LiteLLM? If so, what sets it apart?
  12. 12. kfallah15||context
    Router and model optimization from traffic is the main differentiator
  13. 13. SilenN||context
    Also a hosted marketplace, not just BYOK
  14. 14. ceroxylon||context
    >The gateway adds under 1 ms for BYOK requests

    Amazing! Really brilliant idea, thank you for sharing this project. There is so much ground to cover in the LLM gateway / routing / reporting world, and this is a great start. The Tinker implementation is my favorite part, fine tuning is much better than a sea of context files.

  15. 15. kfallah15||context
    Thanks! We are going to add continual RL via Tinker soon too
  16. 16. akshay_akula||context
    Open source and no markup is the right default for a gateway. The caching question above is the one I would want answered before swapping models though.
  17. 17. SilenN||context
    Ans: we rarely switch, often times it's just a "switch to using this model for your agent"
  18. 18. 0xbadcafebee||context
    You started it a week ago? I look forward to checking back in 3 weeks when you've exited for $1B
  19. 19. SilenN||context
    See you soon
  20. 20. tyre||context
    Looks like first PR is June 24th: https://github.com/experientiallabs/experiential/pull/1

    So, two months. Still impressive!

  21. 21. rdslw||context
    impressive only if using pre-gpt era assumptions about saas/products/software.

    unfortunately a small team can reproduce it in two months, which greatly lowers value of it.

    we, as a collective, have to change our value-judging logic and tune it to post AI world.

  22. 22. SilenN||context
    Thanks for the positivity tyre! If you look at our git history, we pivoted and only started building the gateway recently. Before that we were building research infrastructure that now powers the intelligence features we provide.
  23. 23. swthbht||context
    Very cool. Does your gateway decide effort levels as well? Or just models?
  24. 24. SilenN||context
    Yep! One interesting example is often Opus 5 on low reasoning ~= Opus 5 on high reasoning.
  25. 25. sangwook||context
    What online signal recalibrates simulated rankings against actual task success? Also do you have a plan to support semantic caching at the router level?
  26. 26. kfallah15||context
    For the online signal, we use a LLM judge with a rubric calibrated offline by the user via TUI. UX of the calibration is a major focus area. Semantic caching is interesting, open to supporting it but not currently planned.
  27. 27. forgetme2020||context
    what's the business model here. How does experiential labs make money
  28. 28. kakugawa||context
    They make money on enterprise plans: https://www.experientiallabs.ai/pricing#enterprise

    Look at the Intelligence features in the Enterprise plan:

    * Per-prompt model optimization

    * Caching

    * A model you own, trained on your traffic

  29. 29. kfallah15||context
    yep, it will be through enterprise licenses and our own hosted platform built on the repo
  30. 30. rdslw||context
    the business model, I suspect, is classic rug-pull in some time after building user base.

    proof: boldly claiming being open source in literally first sentence, while cowardly hiding on-by-default telemetry (WTF??) in truly last paragraph of readme.

    sorry to sound harsh, but this is typical old era playbook here.

    in the era of AI, fortunately, such products has much lower value. people and VCs didn’t yet tune to it.

  31. 31. SilenN||context
    Telemetry is off by default. PostHog is for usage analytics on the open source repo. Audit it if you're skeptical.

    We make money off enterprise licenses and hosting models.

    I have strong reason to suspect you either can't read or are a bad actor.

  32. 32. gpiechnik2||context
    great design! i love it
  33. 33. SilenN||context
    Thank you!
  34. 34. nejch||context
    A part of me wishes the open source community would focus making research and industry-backed initiatives like the vLLM Semantic Router rock solid. Then I'd spend less time every month checking if this or that new model router has differentiating over vllm-sr :)

    At least for open source inference, it seems like there's healthy competition centered around vllm/sglang, but 2026 seems to be for model routers what 2025 was for agent harnesses.

  35. 35. d2p||context
    > and use your traffic to (opt in) train you a model.

    Is there more info on this? I'm curious exactly what it is. Is it fine-tuning/LoRA on some base model? Don't cloud providers encrypt reasoning now - does that prevent this?

  36. 36. croemer||context
    Probably shouldn't call it "Open router" in the title as that's a specific brand. Maybe you meant "we built something like OpenRouter".

    Rereading I see you wrote "open OpenRouter", which looks a bit like a typo at first glance.

  37. 37. mongrelion||context
    Congrats on the launching of your product. I will be taking it for a spin to compare it with these other products that seem to be competing directly with what you have to offer:

    - https://github.com/ENTERPILOT/GoModel - https://github.com/maximhq/bifrost - https://github.com/BerriAI/litellm

    Would you care to share what makes experiential different?