
Now loading...
Diogo Almeida, a former OpenAI researcher and key contributor to the development of reinforcement learning from human feedback, expressed deep disappointment in how AI language models have been optimized for human communication rather than practical automation. “We have lightning in a bottle, and yet it isn’t useful,” Almeida remarked. He concluded that while advancements in human language processing have been impressive, they fall short of meeting automation needs, as computers require a different kind of interaction.
In pursuit of a solution, Almeida founded TypeSafe AI after departing OpenAI two years ago. Recently, the startup unveiled a new transformer-based model named Jev. Unlike traditional large language models, Jev focuses on generating probabilities or “calibrated decisions” rather than text output.
This innovative approach has several advantages; it enhances both speed and cost efficiency, while pre-defined user inputs minimize the likelihood of hallucinations. Jev’s model architecture allows for the generation of output tokens at a fraction of the cost associated with conventional models.
Interest in Jev has surged among developers, with its API usage at one point overwhelming TypeSafe’s capacity to serve requests. The model has proven particularly beneficial for software automation. Software developers have noted that Jev provides a more economical and efficient means of embedding intelligence into their applications.
Pranit Sharma, a software engineer at Vercel, shared his experience in which his company replaced OpenAI’s ChatGPT Luna 5.6 with Jev to improve command safety classifications. He reported that using Jev resulted in processing speeds that were five to 18 times faster with superior accuracy.
In another instance, Bryo AI’s CTO Nikhil Mudholkar compared Jev to the Gemini model for classifying business emails. While Gemini showed marginally better accuracy, it was significantly more expensive—10 to 20 times higher. Mudholkar highlighted Jev’s provision of confidence scores as particularly advantageous, stating it offered a real probability that aids in automation.
Beyond merely replacing large language models, Jev may also serve as a supplementary tool, monitoring LLM actions to prevent potential misuse. Almeida advocates for Jev as a practical choice for observing LLM behaviors without incurring excessive costs.
Armin Ronacher, CTO of Earendil, noted that Jev’s model empowers users by shifting the responsibility of interpreting confidence levels to them. This means users can discern which outputs to act upon based on probability scores, fostering informed decision-making.
Furthermore, Ronacher indicated that Jev could prove useful in model routing, offering real-time predictions on whether specific workloads necessitate a certain model, potentially streamlining operations at a lower cost.
The moniker Jev pays homage to William Stanley Jevons, a 19th-century economist known for his paradox concerning commodity usage and costs. Almeida envisions that as the cost of implementing intelligent solutions decreases, their integration throughout various applications will likely expand, akin to the early days of the internet.
Although Almeida remains reticent about Jev’s underlying architecture, some industry observers speculate that it utilizes an open-weight large language model as a foundation. TypeSafe refers to Jev as a “System One model,” aimed at intuition rather than logic, and trained on synthetic data through a novel method termed “reinforcement learning from calibrated decisions.” Almeida emphasized the strategic decision to concentrate on generating synthetic data as a key driver of the company’s success.
As it stands, Jev is unique in the market, but Ronacher forecasts the emergence of competitors as its practical applications become more widely recognized. He remarked that the current ease of using LLMs might have stifled innovation in the sector.
TypeSafe intends to develop additional iterations of its model for various applications. Almeida differentiated his company from traditional frontier labs, expressing a desire to prioritize intelligence as its primary output rather than fostering hype or speculative ambitions common in the industry.
