Jev: New frontier model 40-400x cheaper and 20-200x faster

(typesafe.ai)

183 points | by albelfio 58 minutes ago

32 comments

  • Oras 0 minutes ago
    > Our first public model is Jev, available today in early access. Jev achieves similar levels of intelligence on System One tasks compared to existing LLMs, while being two orders of magnitude faster and more efficient.

    Where is the 20-200x? misleading title.

    Also I don't get it, is that decision tree for automation that can apply to generic problems?

  • mushufasa 19 minutes ago
    I would love for things like this to be accessible via hubs like open router or AWS bedrock. It's hard to justify adding new model vendors directly with all the heightened concerns about privacy and security, but if bold new capabilities are added to a centralized already-vendor like AWS, technical people can adopt them without going through a whole compliance/purchasing/vendor review process. And an extra middleman tax is well worth it when the cost savings of the model itself can be one-two orders of magnitude.
    • cheeze 4 minutes ago
      Isn't openrouter the exact opposite of caring about security and privacy?

      I guess you can choose your provider still? But isn't the point that the lowest bidder is doing inference?

    • oblio 10 minutes ago
      The thing is, in this climate it's hard to believe such tech will remain secret for long.

      So, assuming this is not vaporware, this would raise the tide for everyone because it shows what's possible.

  • ramon156 18 minutes ago
    This sounds good but so far all claims just sound like marketing terms. I'd love to see real proof. e.g. "RLCD" and "parallel sampling" have nothing to back it up.

    also "70-500ms vs 3-329 seconds" are apples-to-oranges unless the LLM baseline is doing comparable work (e.g., long chain-of-thought). If Jev is skipping generation entirely for a narrow structured task, of course it's faster.

    Nonetheless i want this to be true, so I'm looking forward to Jev

    • why_only_15 12 minutes ago
      They have various benchmarks, e.g. how much time it takes them to do wikipedia page -> page games. Jev seems to take the same or fewer hops but in ~10x less time and for ~10x less money.

      It's totally reasonable to compare against LLMs doing chain of thought if it gets comparable performance.

    • vatsachak 10 minutes ago
      It's not an LLM though it's a frontier model on structured data
    • BoorishBears 5 minutes ago
      Did you see the video where it plays Doom, it made it click for me
      • simianwords 1 minute ago
        BTW it was not multi model playing doom, it was passing structured input and getting structured output. Its not what I thought: frames of video passed and real time game play.
  • big_toast 2 minutes ago
    It seems like the docs[0] are a better explanation? The comparison to llm tokens is kinda confusing.

    It looks like the model takes as input a state (structured text? not sure if multi-modal) and a question (as a "Choice", "Score", or "Noul") with some additional augmentations possible. Then outputs the question's answers as appropriate (e.g. a choice, probabilities, confidence).

    On the AI primer page, it looks like they do RLCD from a pre-trained base model?

    [0]:https://docs.typesafe.ai/concepts/system-one

  • dgellow 39 minutes ago
    Side note: it took me more time than I would like to admit to realize that Diogo Almeida isn’t a satirical version of the name Dario Amodei
    • bogzz 7 minutes ago
      That would have to default to Wario Amodei.
    • jakintosh 37 minutes ago
      It wasn't until the demo videos that I realized the post wasn't satirical.
  • jawns 7 minutes ago
    I could see this being fantastic for classification tasks. Last year I shifted from using LLMs for bulk data classification tasks (1M transcripts) to generating embeddings and categorizing based on cosine similarity. It saved a ton of costs and time, but wasn't as accurate as LLMs. This seems like it can give me Terra-level classification ability with the cost/speed I need.
  • bthornbury 2 minutes ago
    Is the tradeoff of the parallel output that we don't get arbitrary string generation? like output # of tokens is fixed ahead of time?

    Either way, really cool and impressive.

  • Gecko4072 2 minutes ago
    For those also confused:

    >Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.

  • skerit 5 minutes ago
    So in theory you could feed it incomplete text, and then ask it for the probabilities of what the next character could be?
    • vatsachak 3 minutes ago
      If you provide it an AST of the english language, yes.
  • vatsachak 25 minutes ago
    It could be used for coding if you gave it an AST.

    If you work at TypeSafe please try this.

    Side note: This is probably how LLMs would perform with better encoders and next-latent prediction, so eventually those will beat this architecture out. Still amazing though.

    • ramon156 17 minutes ago
      I've implemented tree-sitter in pi before, and while it works, I have no real proof it saves me tokens, or is more accurate. I think a better implementation is a model that's trained for AST's, not just "use tool, see what happens".

      I'd love to do research on this when I have the time.

      • vatsachak 11 minutes ago
        Cool project!

        That's what I was insinuating through "better encoder"; the model creating more efficient representations of ASTs using something like JEPA

  • initsecret 17 minutes ago
    > [others] Output tokens: ~5x more expensive than input tokens.

    > [them] Output tokens: FREE (too cheap to meter).

    I'm very confused by this.

    • quotemstr 4 minutes ago
      They're not doing autoregression, so all the outputs are computed in one big forward pass. Very cheap.
  • himata4113 23 minutes ago
    They never show exactly how they use it? Only a bunch of animations of it 'working'. Would like to see the actual code used for the demos!
    • ricardobeat 17 minutes ago
      The doom demo shows the program state / query.
  • albelfio 56 minutes ago
    • magicmicah85 31 minutes ago
      The doom demo is also in the article, for anyone that doesn't want to go to X.com. :)
    • ErneX 36 minutes ago
      • thih9 29 minutes ago
        The doom video is also in the article itself (headline: "Doom").

        I suppose this is the same video as the one from the parent comment, but I don't know for sure - I don't have a twitter account and the above link doesn't work for me.

        • ErneX 3 minutes ago
          I linked to the tweet that has the video because if you are not signed in you cannot see the whole thread of tweets.

          I can see the individual tweets in the browser while not signed in though.

        • yehat 14 minutes ago
          [flagged]
  • moffers 7 minutes ago
    So is it a structured data-based language model? Or is there a model and a harness? Hopefully they’ll open up and explain more.
  • charcircuit 0 minutes ago
    Parallel inference where you don't want a subagent seems niche.
  • bfeynman 20 minutes ago
    Super intrigued by this - large scale automation using LLMs is quite annoying due to deprecation cycles of models from frontier labs and cost of running your own being prohibitive when you have a blend of them.
  • jrickert 34 minutes ago
    Signed up for the beta! :) would love to put this through some real-world shootouts against traditional LLMs to see where this type of model really excels.

    I’m guessing it might be able to replace maybe 40-70% of LLM calls for a given pipeline depending on the business task, cutting the API costs on those calls by an order of magnitude.

  • andai 14 minutes ago
    Why did they pick the name System One? It's not really explained what "System One tasks" and "System One shaped queries" are. Things that need a fast response?

    Does this imply it's a very small model? I couldn't find anything about the model itself.

  • scottyah 38 minutes ago
    Wild that it doesn't generate text. I wonder how its technology compares to Tesla's FSD stack.
  • sim04ful 24 minutes ago
    This sort of stuff almost sends shivers down my spine, it's like i'm looking 5 years into the future.
  • pennomi 31 minutes ago
    > Extraordinary claims require extraordinary evidence so see below for the receipts.

    Yes, that’s the kind of attitude I want to see in these model releases

    • ramon156 16 minutes ago
      But the evidence is not there...
      • pennomi 13 minutes ago
        Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.
        • simianwords 1 minute ago
          They gesture at not using benchmarks for some reason...
  • jceg 9 minutes ago
    > We deliberately chose not to publish performance against public benchmarks. In fact, we plan to only have one-off evals when we make product updates.

    lol, I bet they would publish them if their score on those benchmarks were good.

  • erichocean 24 minutes ago
    I could put this to use today.

    I think we'll see a bunch of different architectures over the next five years.

  • totallygeeky 12 minutes ago
    Woof, that page is hard to read. I don't understand what they've done to the way text is rendering but it's not great for my eyes.
  • whalesalad 17 minutes ago
    What is it about the rendering of this page that is so... off? It almost looks like the entire thing is a <canvas> element.

    edit: looks like a framer export where there is a text stroke being applied :|

  • kypro 1 minute ago
    > Outputs

    > LLMS > Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.

    > Jev > Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores.

    I mean, this isn't even remotely comparable to LLMs so why compare? Also, why are they bringing up AGI given there approach is so restrictive that what they're building literally cannot have the creativity required for AGI? The video is 100% marketing slop...

    The bulk of the application of LLMs is that they generate reasonably reliable text which doesn't need to be defined in advanced. I'm sure there is a niche for this and congrats to the team, but please let's not hype this as if it's the next big thing in AI...

  • larodi 8 minutes ago
    "is this the real thing or is just fantasy"
  • mkrishnan 2 minutes ago
    If this is true, then AI Stock Bubble burst (for Good)
  • hunterbrooks 28 minutes ago
    um what is going on with the outfit changes in the launch video...

    https://x.com/CompleteSkeptic/status/2099925682726002904

  • mkrishnan 2 minutes ago
    If this is true means, AI Stock bubble burst. (For good)
  • esafak 31 minutes ago
    Looks like a great model for NLP.
  • quotemstr 4 minutes ago
    It looks like a specialized encoder-only(-ish) transformer with scalar and ordinal output heads. Acausal in effect, maybe? Probably not even autoregressive?

    I'd use this as a tool an LLM can use for specialized tasks. It's not AI in itself.