Jev plays Doom, disability waits its turn

On 15 September, TypeSafe AI opened Jev in early access, on the very day it announced 40 million dollars raised. Five days later the waiting list was gone; in between, the API had buckled under demand. Developers are delighted, the press ran its customary checks (the "zero hallucination" claim does not hold, the comparisons are home-made), and everyone got to admire the model playing Doom.

Well done. I mean it: a model that decides in under half a second, for next to nothing, with a usable probability, is progress. I have simply seen enough of these announcements to ask the question nobody asked. And for disability, when does it arrive?

What Jev can do

Jev does not write. You give it a "state" (an email, a ticket, a description of a situation) and questions whose answers are fixed in advance. It picks an option from a list, places something on a scale, or says whether a statement is true, with a probability. TypeSafe announces between 70 and 500 milliseconds per request and charges 0.042 dollars per million tokens.

The name comes from Jevons, the economist who noticed that a more efficient steam engine made people burn more coal. The company has turned it into a programme: every drop in the cost of intelligence should open up many more uses. The uses it chose to show are email sorting, customer service, invoices and agent security. For demonstrations, Doom and a link race on Wikipedia. One has the priorities one can.

The uses nobody demonstrated

It does not take a wild imagination to find some on the disability side. Jev's three primitives describe, almost word for word, needs that have been known for decades.

An augmentative communication board offers pictograms or phrases to a person who does not speak, or not fast enough for people to wait for the end of their sentence. Narrowing forty choices down to the six most likely given the conversation is a Choice question. For someone who selects with a switch, cell by cell, every choice removed saves seconds, and a person is waiting opposite. In an invoice pipeline nobody is waiting for anything; yet it is on that kind of pipeline that TypeSafe measured its "193.6 times faster".

Judging whether an administrative text is readable by a person with an intellectual disability, on a scale: Score. The rules of easy-to-read language have existed for a long time, and administrations apply them when they have the budget, which is to say rarely.

Checking that an image description tells a blind student what they need to know, and not just that there is "a diagram with circles and arrows": a true-or-false statement, to run on every document uploaded to a course platform, for a fraction of a cent.

And sorting accommodation requests, which takes weeks in many institutions while the student sits their exams with nothing, would benefit from a machine putting complete files on one side and those waiting for a document on the other.

None of this is an original idea. Everyone has thought of it. It just does not make a demo that goes around social media, and it does not sell to CTOs.

Why it will not come on its own

I will be told that TypeSafe provides the tool and the applications will come from elsewhere. Its own documentation takes care of refuting that. The model accepts only text, and leaves it to others to transcribe everyone who does not communicate in writing. English is its main language; the others are supported, "but not as well". Customers cannot fine-tune it: the same weights serve every account. It is trained only on synthetic data, which the company designed and about which it has detailed nothing. And its calibration, the documentation specifies, is measured across groups of predictions, that is to say on average, where minority populations weigh the least.

An association that wanted to build a communication board in French on Jev could therefore neither teach it its vocabulary nor know whether it has ever seen a conversation resembling those of its users. It can pay per use, like everyone else, and keep its fingers crossed.

The Jevons paradox promises that use explodes when cost falls. It forgets to say which uses. The first served are the ones that pay, and I have yet to see disability at the front of that queue.

The poor relation

TypeSafe invented nothing. In the AI industry, disability occupies the same place as inclusiveness or ecology: brought out when it makes you shine, put away when it costs. In March 2023, for the launch of GPT-4, OpenAI showcased Be My Eyes, an app for blind people, among its partners. The model could describe images, and what could be more moving, to show it, than a blind person having their surroundings described to them. The image went around the world. Access, meanwhile, stayed limited to a few beta testers, and only widened five months later, once the announcement had served its purpose.

With Jev, we do not even get the photo. Disability appears neither in the launch post, nor in the documentation, nor in the page listing the model's nine known weaknesses, nor in the reception as Wikipedia summarises it. It did not make anyone shine this time, so it was not invited.

Nobody is against it, of course. Nobody is ever against it. It is kept for later, for version 2, for the day a customer asks for it with a purchase order, or for the next inclusion campaign, when a picture will be needed.

The next demonstration

TypeSafe showed that Jev could play Doom in real time, from a structured description of the game state. People spent time, in a company that had just raised 40 million dollars, getting a classifier to kill demons in 70 milliseconds.

The same speed, serving a person who speaks with a switch, would take one more dataset, French handled a little better, and published figures on the users concerned. It is not out of reach. It would just have to come before Doom on the list, and I am not holding my breath.

References