What it means to build judgment into software

By Mirco Neri

Most software that calls itself intelligent is good at two things: storing what you give it and generating text about it. Neither of those is judgment. Judgment is deciding what to leave out, and it is the thing I spent most of this year trying to build, mostly by discovering what it is not.

Everything is not equally alive

The first version of Athena's morning briefing was a language model with a lot of context and an instruction to tell me what mattered. It read beautifully. It was also wrong in a way that took me weeks to name: it treated everything I had ever told it as equally alive. A task from three weeks ago and a thought from last night sat side by side with the same weight, because to the model they were both just text. Humans do not work that way. Something you said in June and never mentioned again has a different status from something you said yesterday, and that difference is most of what judgment is.

So the first thing Athena has to judge is time. A view about tomorrow's market should be gone by the end of the week. A rule written after a loss should never be gone. A plan for a quarter is alive until the quarter ends and then becomes history. When you ask Athena what matters today, the things whose time has passed are simply absent. That sounds like housekeeping, and it turned out to be the biggest single improvement the briefing ever had, because judgment starts with the pile being the right size.

Connections are questions

The second thing was harder. A language model asked to find what matters will find connections, because finding connections is what it is for. Early on, Athena told me that a loss on one day was related to a decision I had made the week before. It was fluent, it was plausible, and it was invented. The two records had nothing to do with each other. A human assistant who did that would be fired, and rightly, because an assistant who connects things that are not connected is worse than one who connects nothing.

The rule that came out of that is now the most important sentence in the product: connections are questions, never claims. Athena may notice that two things might be related and ask me. It may not tell me they are. One word of grammar, and it is the whole difference between a tool that thinks with you and a tool that thinks for you and gets it wrong.

Knowing when not to speak

The third judgment is about silence. On a Saturday, Athena does not put work tasks in front of me, whatever the deadlines say, because I told it once that weekends are not for that and it holds me to it harder than I hold myself. When I say a thread is finished, it stops asking about it and lets the memories that depended on it retire, instead of leaving them to compete with what is true now. When a question would be a bad question, because I already answered it or because the thing it is about has moved on, it asks nothing. A briefing that is shorter than yesterday's is often a better one.

What is left for the machine

What is left, once time and restraint have done their work, is the part that actually needs intelligence: reading the shape of a week and saying, in plain language, why today matters and which of the things I said were important has quietly stopped getting attention. That is a real task and a model does it well. But it does it on a pile that has already been sized, and it is forbidden from asserting anything the record does not support.

One question a day

The last piece is the one people find strangest. Athena learns from one question a night. It asks something about how the day went, chosen from what is open and unresolved, and I answer in a sentence or two. That answer is the whole lesson. I used to think an assistant needed to see everything to be useful. I now think the opposite: the less it sees, the more each thing it does see has to earn its place, and that constraint is what forces judgment rather than coverage.

None of this is clever machine learning. It is a series of decisions about what a system should refuse to do. Judgment, in software as in people, turns out to be mostly restraint.

Mirco Neri runs a commodity trading firm in Dubai and built Athena. athenaos.net

How to connect Claude to a memory that is actually yours