<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>geord.ee</title>
    <description>Personal Weblog</description>
    <link>https://geord.ee/</link>
    <atom:link href="https://geord.ee/feed.xml" rel="self" type="application/rss+xml" />
    <pubDate>Sat, 03 Oct 2026 08:54:48 +0100</pubDate>
    <lastBuildDate>Sat, 03 Oct 2026 08:54:48 +0100</lastBuildDate>
    <generator>Jekyll v4.4.1</generator>
    
      <item>
        <title>Let the Model Synthesise</title>
        <description>&lt;p&gt;For the last few months I have been working on a project to classify requests from our store colleagues. The classification has 150+ classes. With a small set of classes it is relatively easy to get it right. Sentiment analysis, for example, is positive, negative or neutral. With 150 classes there are 150 ways things can go wrong, or even more.&lt;/p&gt;

&lt;p&gt;I started with what seemed the obvious approach. I gave the model the user’s message and all classes with their descriptions in a single prompt, and asked it to pick one. Zero-shot classification. It was a gamble, and it did not pay off. Accuracy was around 18% - too low to do anything practical with.&lt;/p&gt;

&lt;p&gt;The next attempt was hierarchical classification. I split the 150+ classes into a three-level hierarchy. The model would start with around 30 categories at the top level, pick one, and then go down a level, and then another. That took us to around 30%, and not much further. The reason was structural. Once the model took a wrong turn at the first or second level, there was no way back to the correct class. The hierarchy had committed to the wrong branch ahead of time.&lt;/p&gt;

&lt;p&gt;So I added context. Past conversations - messages that had already been classified - were used along with the class descriptions. They reinforced the right behaviour, and in some cases were allowed to override the hierarchical classification altogether. Accuracy went up to 40%, and with further tuning, to 60%. A significant improvement, but still short of what we needed for production.&lt;/p&gt;

&lt;p&gt;At this point we stepped back and noticed something a bit funny. We were using an LLM to traverse a hierarchy, a graph. The hierarchy is a tree, and two or three prompts were walking it in one direction, from the root to a leaf. But a language model is not the obvious tool for walking a graph.&lt;/p&gt;

&lt;p&gt;So we did something completely different. We walked the graph using confidence, similarity and a confusion matrix, with tried and trusted embedding models. We also added new edges between the leaves - “often misclassified as” and “very similar to”. When the traversal reaches a leaf, these edges let it jump across to a leaf in another branch altogether. The tree became a graph. A wrong turn early on was no longer necessarily final.&lt;/p&gt;

&lt;p&gt;And the best part is that the traversal did not need an LLM anymore. The graph and the similarity models identified a handful of potential leaves. We collected those candidates together, along with the recent context from the conversation history, and handed them to the LLM. The model now had a small and rich context to reason over, instead of 150+ options or a single path down a tree. I call that final prompt the synthesis prompt. Accuracy went up to 80%.&lt;/p&gt;

&lt;p&gt;The solution is now a hybrid pipeline - embedding models and a graph narrow down the candidates, and a large language model makes the final call. The LLM alone could not get us there.&lt;/p&gt;

&lt;p&gt;We often choose one technology and pin ourselves down to it. Today that technology is usually the LLM, and the instinct is to solve every part of the problem with a better prompt. But that is not how we operate as human beings. We draw on a continuum of things we have learned from a wide range of experience - a bit of structure, a bit of logic, a bit of reasoning, and hard evidence that comes from statistics and past observation. Then we synthesise.&lt;/p&gt;

&lt;p&gt;I believe the solutions built on LLMs should work the same way. Let each technique do what it is good at, and let the language model do the synthesis. In my opinion that synergy - between the old and the new, the classical and the generative - is what really makes LLM-based solutions powerful.&lt;/p&gt;
</description>
        <pubDate>Sun, 13 Sep 2026 23:41:22 +0100</pubDate>
        <link>https://geord.ee/2026/09/let-the-model-synthesise</link>
        <guid isPermaLink="true">https://geord.ee/2026/09/let-the-model-synthesise</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>Ask for the Program, Not the Answer</title>
        <description>&lt;p&gt;A couple of weeks ago I worked with our business team on applying generative AI to a large dataset - around 130,000 rows in a spreadsheet. The model kept trying to retrieve and summarise. That appears to be its default behaviour. The team had tried everything reasonable. They split the data into smaller chunks. They wrote step-by-step instructions to walk the model through the retrieval. The answers kept drifting.&lt;/p&gt;

&lt;p&gt;Then it struck us. The model was searching, not filtering or aggregating. Search is typically an approximate, probabilistic operation. What we needed was deterministic - the kind of result that comes out of a SQL query or a Python script. So we stopped asking the model for the answer and asked it to write a program that would filter, aggregate and project the data. And it worked. The change in prompt was little more than “use the code interpreter to find the answer”.&lt;/p&gt;

&lt;p&gt;This should not have surprised us. Generative AI is built on large language models. By definition they are good at language, not numbers. But numbers can be operated on through a programming language, and a programming language is language. The model does not need to crunch the data; it needs to write the thing that crunches the data. It generated Python code, ran it against the sheet, and returned the right result every time.&lt;/p&gt;

&lt;p&gt;A few weeks before that I had been working with Model Context Protocol - the protocol that injects context into a model invocation. It was a proof of concept, and we had wrapped an API with an MCP server. At runtime the model called the tool and the entire API response landed in the context window. It was a large response. The context bloated, the costs added up with every run, and it was slow.&lt;/p&gt;

&lt;p&gt;So we changed the approach. The model still decided which tool to call, but instead of calling it directly, it wrote a program that called the tool, found the answer, and injected only that answer into the context. Speed improved. Costs came back under control. Note that we were using MCP as an API, which is what it had wrapped in the first place. MCP became a discovery mechanism rather than a data pipe.&lt;/p&gt;

&lt;p&gt;This is probably not the way to build MCP servers. It is fairly well understood by now that wrapping an API in MCP is neither effective nor efficient. But the workaround revealed the same pattern as the spreadsheet. Give a model a code interpreter and it will solve a problem in a completely different way than it would otherwise.&lt;/p&gt;

&lt;p&gt;We often compare Generative AI to a brain. Tools give it senses and limbs. The code interpreter is the simplest way for a model to act on the environment it finds itself in, or at least on the sandbox it has been given. My view is that code interpretation will become an essential feature of generative AI solutions rather than an optional one. In human terms it is the equivalent of learning the multiplication table, or a formula, or a skill - the thing you reach for so you do not have to reason from first principles every time.&lt;/p&gt;

&lt;p&gt;So the next time you have a problem involving structured data or hard logic, do not ask the model for the answer. Ask it for the program.&lt;/p&gt;
</description>
        <pubDate>Tue, 08 Sep 2026 09:13:12 +0100</pubDate>
        <link>https://geord.ee/2026/09/ask-for-the-program</link>
        <guid isPermaLink="true">https://geord.ee/2026/09/ask-for-the-program</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>The Mythical Agent-Month</title>
        <description>&lt;p&gt;The Mythical Man-Month by Frederick P. Brooks Jr. is a collection of essays on software engineering, first published in 1975. Though it is popularly known for its main theme of adding manpower to an already delayed project, the book is not about project management. It is about the balance between design and effort, between labour and thought.&lt;/p&gt;

&lt;p&gt;In the age of AI agents writing code, we may rethink how the essays and concepts in the book apply in an organisation of agents. Here are a few essays with the original idea and a reframing in the context of agent organisation.&lt;/p&gt;

&lt;h2 id=&quot;the-tar-pit&quot;&gt;The Tar Pit&lt;/h2&gt;
&lt;p&gt;It is easy to build an MVP. But to scale that up into a system or a product, it may take up to 9x the effort or cost, due to testing, integration, deployment and documentation at scale.&lt;/p&gt;

&lt;p&gt;An agent organisation cuts the effort and potentially the costs significantly. However, if we are not careful, it ends up generating a lot of half-owned software nobody understands. The tar is no longer the effort, but the accumulated intent nobody owns.&lt;/p&gt;

&lt;h2 id=&quot;the-mythical-man-month&quot;&gt;The Mythical Man-Month&lt;/h2&gt;
&lt;p&gt;Adding more people to an already late project delays it further, due to tribal knowledge and communication overheads.&lt;/p&gt;

&lt;p&gt;In an agent organisation the ramp-up time almost vanishes. We may be able to add more agents with low overheads on communication and training. But the underlying constraint may not be communication. It may be about building shared knowledge and understanding. In the world of agents, if the shared knowledge is not factored in, the outputs diverge, and the project is delayed.&lt;/p&gt;

&lt;h2 id=&quot;the-surgical-team&quot;&gt;The Surgical Team&lt;/h2&gt;
&lt;p&gt;Structure large teams like a surgical unit, with a chief programmer/architect, and the rest of the team supporting with different specialisations.&lt;/p&gt;

&lt;p&gt;This is one of the most visible aspects of an agent organisation. Senior engineers are able to command a host of agents effectively. The coding agents are multi-specialised, and may contain various aspects of a surgical team within one agent’s scope. Also note that the human chief programmer may not be able to cope with the speed at which AI agents write code. From my practical experience, I was able to drive up to five coding agents before it became too overwhelming.&lt;/p&gt;

&lt;h2 id=&quot;the-second-system-effect&quot;&gt;The Second-System Effect&lt;/h2&gt;
&lt;p&gt;An architect’s first system is usually clean and simple, while their second system is often dangerously over-engineered, stuffing it with every idea they suppressed in the first.&lt;/p&gt;

&lt;p&gt;In an agent organisation this is a real risk, and a risk that will materialise at machine speed. Today’s agents will implement any request and gold-plate anything. And the programmers are often happier to add more capabilities because now they can. Restraint in both human instructions and agentic process is probably the solution. Of late, I see Claude pushing me back on certain complex requests. I wonder whether that’s part of the system instructions.&lt;/p&gt;

&lt;h2 id=&quot;plan-to-throw-one-away&quot;&gt;Plan to Throw One Away&lt;/h2&gt;
&lt;p&gt;The first version of the system becomes non-viable quickly, since users will change requirements once they see the working product. One should intentionally plan the first system to be a prototype, which can be thrown away without much regret.&lt;/p&gt;

&lt;p&gt;In an agent organisation, we don’t just throw away the first prototype. We could throw away ten or twenty prototypes, and some built concurrently. Or it could be an evolutionary search, finding the fittest design among different mutations. Brooks almost retracted the idea in favour of incremental development. He did not need to. We can plan to throw ten away.&lt;/p&gt;
</description>
        <pubDate>Thu, 02 Jul 2026 00:42:30 +0100</pubDate>
        <link>https://geord.ee/2026/07/the-mythical-agent-month</link>
        <guid isPermaLink="true">https://geord.ee/2026/07/the-mythical-agent-month</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>Printing Software</title>
        <description>&lt;p&gt;I often write in analogies… how the new AI wave compares to the dotcom boom or the industrial revolution. There is something to be learned from every transformation, whether it is from the factory floor or the financial markets. When we remove the specifics of the technology advancements, it becomes an economic phenomenon or a supply chain problem.&lt;/p&gt;

&lt;p&gt;Looking at the way software is written by AI, I cannot help but compare it to 3D-printing in the manufacturing domain. 3D-printing was supposed to revolutionise manufacturing, but as of today it is still a rounding error in the overall manufacturing sector. It is worth noting that the technology is improving at a fast pace, and the adoption is also increasing. Where it makes the biggest difference is prototyping. Watch the popular maker channels on YouTube: they use 3D-printing to prototype and validate designs before committing to more expensive materials and scaled-up production. The adoption in the final production process is also increasing, but not at the scale or pace to replace the current techniques or processes.&lt;/p&gt;

&lt;p&gt;AI-assisted coding is also going through such a stage. Prototyping software is easy. It is instantaneous. Using popular vibe-coding platforms, one does not need extensive knowledge of software engineering or expensive tooling. It runs in the cloud, and most of the time auto-deploys. But it is still a prototype. It might break when you add different categories of users or scale to production workloads.&lt;/p&gt;

&lt;p&gt;There are a few reasons that stand in the way of 3D-printing adoption. They include consistency in quality, overheads during post-processing, the economics of scaling, speed and throughput for scaled-up volumes, and limitations in the materials suitable for 3D-printing.&lt;/p&gt;

&lt;p&gt;I believe there are similar limits in software. The quality of the software generated by AI varies widely based on the people who prompted it and the technical vocabulary they used. There are of course overheads after the code is generated, with increasing demand for people to clean up AI-generated code. And AI-generated code has certain recognisable patterns, even when AI is used to generate the designs, which limits the variety and creativity.&lt;/p&gt;

&lt;p&gt;The technological advancements in 3D-printing and code-generation are real and useful. In the short term it makes sense to leverage these nascent technologies to build prototypes, familiarise ourselves with them, push boundaries and establish best practices. And in some cases, take them to production and scale up. 3D-printing promises to cut the waste of subtractive manufacturing through additive techniques, and even create tailored products on-demand. The promises of AI-assisted software engineering are quite similar. They might be on the same trajectory. But software trends do move faster than materials and manufacturing.&lt;/p&gt;
</description>
        <pubDate>Wed, 01 Jul 2026 00:45:42 +0100</pubDate>
        <link>https://geord.ee/2026/07/printing-software</link>
        <guid isPermaLink="true">https://geord.ee/2026/07/printing-software</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>Credentials for AI Agents</title>
        <description>&lt;p&gt;When we employ people for a job, we verify their credentials - their educational qualifications, certifications and work experiences. In the world of AI agents, similar to people, AI agents will need credentials. Human beings acquire theirs through the education system. We are trained for fifteen to twenty years, and that training qualifies us for certain jobs. The proof is in the certificates issued by universities and the experience letters from the companies where we worked.&lt;/p&gt;

&lt;p&gt;AI is recent, and agents are more recent still. So how do we make sure they hold the credentials to perform critical tasks in the enterprise? We would want to know that an agent has undergone rigorous training in the domain it operates in, and that it has been vetted by a body qualified enough to vouch that it is capable of the work. Today we have none of that. Is it enough to take the claims of the companies that build these models at face value? Is it enough to look at the benchmarks? Is it enough if the AI experts inside each of these companies declares it so? The marketing claims are accepted because there is nothing else to accept.&lt;/p&gt;

&lt;p&gt;Part of the difficulty is that today’s models are general-purpose. We use the same model for nearly everything, and we shape it for a task at the point of deployment, through a system prompt, a set of tools, some context, and a temperature setting. The competence does not really sit in the model; it sits in that arrangement or architecture around it. It is hard to certify something that is reassembled differently every time it is deployed.&lt;/p&gt;

&lt;p&gt;This is why I expect the future to be different. We will have models post-trained for standard duties, specialised for an industry, a technology domain or a functional domain. When the competence is trained into the weights rather than patched on at deployment, the thing being certified becomes durable. A certifying body can then examine such a model and issue a verifiable credential - a tamper-evident attestation, signed by the issuer, that anyone can check without having to trust whoever presents it. And the credential can bind to something concrete: the model’s name and the hash of its weights.&lt;/p&gt;

&lt;p&gt;You cannot cryptographically verify that the doctor in front of you is the same person the medical board examined years ago. But you can verify that the model answering your query is the one that was certified. And by the way, the model does not have a bad day. It does not forget its training, and it does not quietly drift. Frozen weights are a more honest subject of certification than a human ever is.&lt;/p&gt;

&lt;p&gt;The competence does not end with the weights, though. A domain model is only as good as the literature, the knowledge bases and the regulations it was trained on, so the credential has to carry the provenance and the currency of that knowledge too. A credential should not say “certified for medical device regulation”. It should say “certified against MDR 2017/745”. That is exactly how we scope human professionals. A lawyer is licensed in a jurisdiction, as of the law that exists today, with an obligation to keep current. The credential carries the boundary of the knowledge, not merely its existence. It also explains the expiry. The credential lapses not on an arbitrary date, but on the date the corpus goes stale.&lt;/p&gt;

&lt;p&gt;The same credential has a second use, beyond proving capability. It gives the agent an identity and a key, and that lets it sign what it does. Once agents act independently, the provenance of their decisions matters as much as their qualifications. Today an AI decision is logged into a database that can be edited, deleted or quietly rewritten. If instead each decision is signed by the agent’s credential, every entry carries proof of which certified agent made it, and that it has not been altered since. The immutable ledger keeps the record in order and tamper-evident; the signature makes each entry attributable and authentic. The more interesting version is a record co-signed by the parties affected by a decision, so that it is shared and verifiable rather than owned by one side.&lt;/p&gt;

&lt;p&gt;We have all but forgotten about blockchain in the rush towards AI. But it is worth a second look. It was built to prove provenance and authenticity, and that is precisely what a credentialed, accountable agent needs. It is interesting how a technology returns to play a role it was not invented for. Blockchain arrived early, found its moment fade, and is available again just when AI needs a way to prove that an agent is who it claims to be, trained for what it claims to do, and accountable for what it has done.&lt;/p&gt;
</description>
        <pubDate>Tue, 30 Jun 2026 00:48:27 +0100</pubDate>
        <link>https://geord.ee/2026/06/credentials-for-ai-agents</link>
        <guid isPermaLink="true">https://geord.ee/2026/06/credentials-for-ai-agents</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>These are a few of my favorite editors</title>
        <description>&lt;p&gt;Three decades, twelve editors. From vi in 1995 to Helix in 2026. The languages, jobs, and habits each one carried. A nostalgic look back, written before editors became viewers.&lt;/p&gt;

&lt;hr /&gt;

&lt;h5 id=&quot;vi-1995&quot;&gt;vi. (1995)&lt;/h5&gt;
&lt;p&gt;That’s how it started. We had Wipro Unix server with dumb terminals. Writing mostly C programs for hobby.
&lt;em&gt;C, Fortran&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;turbo-c-editor-1998&quot;&gt;Turbo C Editor. (1998)&lt;/h5&gt;
&lt;p&gt;Not a real favorite. It had keyboard shortcuts and mouse support though. It took me through final year project to design concrete structures based on structural analysis outputs, constraints and assumptions.
&lt;em&gt;C&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;tedit-1999&quot;&gt;Tedit. (1999)&lt;/h5&gt;
&lt;p&gt;Terminal editor of Tandem computers. T6530 emulation. We had more function keys than our open systems colleagues. And with OutsideView we were cooler than our mainframe colleagues!
&lt;em&gt;Screen COBOL, TAL, TACL&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;codewright-2001&quot;&gt;CodeWright. (2001)&lt;/h5&gt;
&lt;p&gt;This was probably the best editor I have used. Syntax highlighting, customisation, macros. It was quite ahead of its time. I was quite sad to see it wane away into oblivion.
&lt;em&gt;TAL, TACL, C89&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;notepad-2004&quot;&gt;Notepad++. (2004)&lt;/h5&gt;
&lt;p&gt;New place. No access to CodeWright. Started with Notepad++. I could define the syntax highlighting for TAL and TACL using user-defined languages. Plus it did not need installation, and that was quite convenient. It was with me for a good 9 years.
&lt;em&gt;C++, TACL, SQL, Perl, PHP&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;gedit-2006&quot;&gt;gedit. (2006)&lt;/h5&gt;
&lt;p&gt;I installed Ubuntu on my laptop. Breezy Badger. There was vim, but I was more drawn to gedit. Not much of programming. More of getting around in the Linux world.
&lt;em&gt;Python, Shell Scripts&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;textmate-2011&quot;&gt;TextMate. (2011)&lt;/h5&gt;
&lt;p&gt;Got my first MacBook Pro. And TextMate. It was underused for a couple of years, until I got into Ruby on Rails. And then TextMate was the thing. Trying to copy Ryan Bates from Railscasts. Another editor I missed.
&lt;em&gt;Ruby, CoffeeScript, HTML/CSS&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;sublime-text-2012&quot;&gt;Sublime Text. (2012)&lt;/h5&gt;
&lt;p&gt;I got it in 2012, but it was my second editor for a long time. And somewhere along the way, Sublime Text overtook. I think it was the Cmd+P shortcut in version 2 that made the difference. And the regex support. I still use Sublime Text (and Sublime Merge). The longest running editor in my toolbox.
&lt;em&gt;Python, Ruby, JavaScript, HTML/CSS, Dart, Perl, PL/SQL, Lua, Shell Scripts&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;vim-2014&quot;&gt;vim. (2014)&lt;/h5&gt;
&lt;p&gt;I made my first return to commandline. Vim and bash were mandated to everyone in my startup. Vim gyms and extensions. The startup didn’t last long. But vim stayed in the background. Extensions vanished, but with servers in the cloud vim remains handy.
&lt;em&gt;Ruby, Python&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;vs-code-2018&quot;&gt;VS Code. (2018)&lt;/h5&gt;
&lt;p&gt;I had tried different Electron-based editors from the time of Atom. I never liked any of those, including VS Code. But then I started working with Java and Spring Boot. And VS Code’s support for Java was quite good, mostly for microservices. I did not want Eclipse or IntelliJ. VS Code became the dominant editor with the ecosystem of language servers and extensions. Until it started getting bloated. And then I removed it.
&lt;em&gt;Java, Python, JavaScript, TypeScript, Go, C#, HTML/CSS&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;zed-2025&quot;&gt;Zed. (2025)&lt;/h5&gt;
&lt;p&gt;I am writing this in Zed. After I removed VS Code, Zed took its place. All the AI features are disabled. It’s a good editor. It has not replaced Sublime Text completely though.
&lt;em&gt;Go, TypeScript, Python&lt;/em&gt;&lt;/p&gt;

&lt;h5 id=&quot;helix-2026&quot;&gt;Helix. (2026)&lt;/h5&gt;
&lt;p&gt;New kid on the block. My latest attempt to return to commandline. If it was not for the vi/vim muscle memory, it would be the primary editor. It’s the default editor in the command line.&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;Postscript: Of late AI-assisted engineering has greatly affected how I use editors. I wrote this list to remember the journey. The years when I typed every character. Before the editors become viewers.&lt;/p&gt;
</description>
        <pubDate>Mon, 29 Jun 2026 00:11:18 +0100</pubDate>
        <link>https://geord.ee/2026/06/favorite-editors</link>
        <guid isPermaLink="true">https://geord.ee/2026/06/favorite-editors</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>Job Descriptions for AI Agents</title>
        <description>&lt;p&gt;There is a pattern in AI user interfaces, especially for the chat-based conversational interfaces. User asks a question or explains the problem. AI responds with a good-enough answer. User then follows up with more relevant information until an acceptable answer or solution is achieved. This suits the probabilistic nature of AI. This was the most common pattern in retrieval-augmented generation (RAG) architectures.&lt;/p&gt;

&lt;p&gt;Contrast this with a human-to-human conversation where the first question will be responded with a counter-question or request for clarification, instead of a good-enough answer. As models became more capable and tokens became cheaper, we have introduced reasoning models and newer patterns. The new architecture is known as ReAct - reasoning and acting. AI agents operate a loop of thinking, acting and observing before responding to the user. These answers tend to be more definitive as the agentic loop does the follow-up without user interaction.&lt;/p&gt;

&lt;p&gt;“Thinking tokens” have dramatically increased token consumption, and oftentimes users have less control over the loop, except for some high-level settings such as low, medium, high and ultra. We can influence thinking characteristics using the system instructions or agent’s personality definition as well.&lt;/p&gt;

&lt;p&gt;In general the agentic capabilities are made affordable using faster inference and cheaper tokens. In my opinion a better context is more efficient and cost-effective than larger models and higher thinking effort. However, better context is often too generic or vague to explain. The right context in software engineering is quite different from that in customer support.&lt;/p&gt;

&lt;p&gt;So, here is a proposal. We should start considering AI agents as assistants or employees. We should assign them organisational roles, give them a personality, a job description, and training manuals. If we give AI agents enough information to answer questions or carry out tasks, they are more likely to succeed.&lt;/p&gt;

&lt;p&gt;I have tried defining the personality and job description of a model with mixed results. Today we use the same model (Gemini 3.x or Claude 4.x or GPT 5.x) for most of the tasks. I consider those general-purpose models. In the real world, when we employ someone, we look for their pre-training, their area of specialisation during education. Whether they have a bachelor’s degree or a master’s degree. Whether they have relevant industry experience. In the future we will have AI models that are specialised in a domain to be employed. That makes the job descriptions more compact, and easier to explain using industry or domain vocabulary.&lt;/p&gt;

&lt;p&gt;The second aspect is the tribal knowledge within the organisation. AI agents do not socialise near the water cooler, in the cafeteria, or in a quick chat across the desk. They will lack the informal, cultural and tribal knowledge of an organisation unless it is formalised. It’s a difficult task and it requires a deliberate change in culture and process to let AI know the details, changes and exceptions of the context. And it may be a short-term challenge, as AI automates more tasks and processes it will learn or retain such details.&lt;/p&gt;

&lt;p&gt;The organisational role and job descriptions help in another way too. It helps to set, track and evaluate the cost and performance of an agent. Hopefully it will be much easier to scale up/down or decommission an agent depending on the cost and performance metrics.&lt;/p&gt;

&lt;p&gt;Recently, when I deployed an agentic solution for a business process, the business leader insisted that we should give a name to the agent. We haven’t yet, but I think it is a good idea, considering the fact that the agent operates semi-autonomously among the people. What I have not considered in this article is a completely different operating model with AI agents in the picture. Even in such a case, the co-existence of AI agents with humans in the organisational hierarchy is probably a good transition phase.&lt;/p&gt;
</description>
        <pubDate>Sat, 27 Jun 2026 23:06:10 +0100</pubDate>
        <link>https://geord.ee/2026/06/job-description-for-ai-agents</link>
        <guid isPermaLink="true">https://geord.ee/2026/06/job-description-for-ai-agents</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>Data Labeling is the New Data Mining</title>
        <description>&lt;p&gt;In 2011 I worked on Oracle Spend Classification, a data mining project. Machine learning algorithms classified products and parts into various spend categories for better control over the spend, contracts and even the design. Unlike the previous efforts the machine learning ran autonomously with a great degree of accuracy. But it worked on structured data, the product master, purchase orders, goods receipt and so on.&lt;/p&gt;

&lt;p&gt;However, the world is mostly unstructured. The things we see, the things we do, the words we speak, the music we play. Computers are remarkably inefficient at making sense of all that. We have to annotate and label the data for computers to comprehend. And so, in the age of large language models, the term data mining takes on a new meaning. The real mining is no longer done by an algorithm. It is done by people.&lt;/p&gt;

&lt;p&gt;Today the Financial Times has published a film on data labeling. People in rural villages strap a GoPro or smartphone to their heads and go about their day - stitching clothes, folding laundry, washing dishes. Everyday human activity, captured as data. Elsewhere, another set of people annotate those frames. This is a car. This is a bag. We send miners down a pit to dig for coal and gold. Today we send people down into their own lives to dig for data. That is the irony few notice: data labeling is the real data mining.&lt;/p&gt;

&lt;p&gt;The supply chain of AI is already global and entirely virtual. The raw material is scattered across the internet. It is mined where labour is cheap, refined where capital and computing power are concentrated, and consumed everywhere. Like petroleum, mined data is refined through training into models, and the finished goods carry far more power and price than the raw materials ever did.&lt;/p&gt;

&lt;p&gt;This is where the conversation about AI sovereignty feels incomplete. It assumes the raw material is freely available and that mining (labeling) may be a solved problem, so that refining becomes the main challenge. I am not sure that holds. The resource and the mining have to be solved in their own context too. A nation cannot build a model that understands its people if the data and the labeling for that context do not exist.&lt;/p&gt;

&lt;p&gt;This is also why smaller, context-aware models may lead the way. They do not demand the same impossible scale.&lt;/p&gt;

&lt;p&gt;There is a trigger to all this. Governments have begun to signal who may access the most capable models. Once a government starts controlling a resource (such as a refined, state-of-the-art model) others are forced to consider their own.&lt;/p&gt;

&lt;p&gt;We are watching a global supply chain take shape in real time, with all the familiar tensions of resource, refinement, and sovereignty - only this time the raw material is us, our words, our actions and our lives.&lt;/p&gt;
</description>
        <pubDate>Fri, 26 Jun 2026 20:42:00 +0100</pubDate>
        <link>https://geord.ee/2026/06/data-labeling-is-the-new-data-mining</link>
        <guid isPermaLink="true">https://geord.ee/2026/06/data-labeling-is-the-new-data-mining</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>AI and Overproduction</title>
        <description>&lt;p&gt;In February I decided to move to the command line. I started assembling the tools to live primarily in command-line mode. One of the tools I missed badly was a good Markdown reader. I searched far and wide. There were not many good options. Glow was the most convincing choice, but I was hitting bugs and limitations. So I decided to build one for myself. Armed with Claude Code, I built one in a couple of weeks - a command-line Markdown viewer with a simple editor and readability metrics. I shared it with some friends and decided to let it cool down before sharing it widely.&lt;/p&gt;

&lt;p&gt;Then, a couple of months later, I started observing something interesting. There was a wave of Markdown viewers and editors. I saw new Markdown tools when I ran brew update. I saw announcements in my Hacker News and Reddit feeds. Of course, it was not just command-line tools; a lot of them were GUI apps. Markdown tools are a low-barrier entry point for tinkering with AI code generation. So, have I given up my Markdown editor? No, not yet. I have tested a few, some carefully crafted, some completely vibe-coded. Not many of them have the ergonomics and the speed I want. So I am still sticking with my creation.&lt;/p&gt;

&lt;p&gt;AI is leading us to an (or another?) era of mass-produced software, built in bedrooms, dorms and coffee shops. In my opinion, we already have an overproduction problem with consumer goods: cheap parts, fragile assembly, generic interfaces, and keysmash branding. If we are not careful, we will have digital products with cheap parts, fragile integrations and generic interfaces flooding our app stores.&lt;/p&gt;

&lt;p&gt;The long tail is a known phenomenon in the digital world - whether in digital products or digital distribution. AI is lengthening the long tail, at the end of which are single-user systems and single-use systems. It is okay to build DIY tools that stay within one’s environment. In a way, spreadsheets were a platform for DIY tools. What makes it difficult is the noise it generates in the public market. It is acceptable for AI-generated tools to compete in the market and replace existing ones that are not up to the mark, and a fair market does not and should not limit anyone’s entry. However, at what cost? The distribution ecosystems, such as app stores and GitHub, and even human attention are getting strained by the number of new apps and tools. And over time, we will have a lot of abandonware, creating an invisible-but-real software wasteland.&lt;/p&gt;

&lt;p&gt;We will learn to improve the quality and reliability, using libraries, agentic frameworks and newer design patterns that we are yet to define. That might fix the abandonware problem, and we may very well move into another cycle of build-vs-buy in which the build might dominate. However, the core issue of overproduction may still remain.&lt;/p&gt;

&lt;p&gt;In the 19th century, the world faced a similar situation during the Industrial Revolution - for example, when mechanised cotton spinning made textiles quite cheap. Foreign markets were an option at that time. Today AI is a global phenomenon. Then followed a series of crises and price deflations, cartels and consolidation. Eventually the world brought supply and demand back into balance, through consolidation of competitors and supply generating its own demand.&lt;/p&gt;

&lt;p&gt;Will AI make software cheaper? If so, will enterprises consume more software by building custom solutions than by buying off-the-shelf software? Is overproduction of goods and tools necessary? Is there an anarchy of production in software? Are the current costs of AI artificially low and are they incentivising overproduction? We have a lot to question, think about and answer in the coming days.&lt;/p&gt;
</description>
        <pubDate>Fri, 26 Jun 2026 00:02:16 +0100</pubDate>
        <link>https://geord.ee/2026/06/ai-and-overproduction</link>
        <guid isPermaLink="true">https://geord.ee/2026/06/ai-and-overproduction</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
      <item>
        <title>New Wine in Old Wineskins</title>
        <description>&lt;p&gt;There is a recurring debate about AI replacing human activities - in coding, design, shopping, negotiation, and medical diagnosis. What we often overlook is that we are asking an inherently probabilistic system to perform tasks that demand determinism and precision.&lt;/p&gt;

&lt;p&gt;My view is that we need to redefine the processes themselves.&lt;/p&gt;

&lt;p&gt;Let us consider the process of design, whether it be web, software, interior, or industrial. Traditionally, we begin with requirements and exact measurements, translating them into prototypes, diagrams, or CAD models. By the time a design is visualised, significant effort has already been invested. Every iteration from that point is costly.&lt;/p&gt;

&lt;p&gt;AI can generate those designs almost instantly. But it risks imprecision, and hallucinations are a real concern.&lt;/p&gt;

&lt;p&gt;What if we redefine the process to include an inspiration step, powered by AI? It accelerates early iterations to a point of acceptable direction, which can then be refined using conventional design tools. AI as a starting point, not a finishing line.&lt;/p&gt;

&lt;p&gt;The same thinking applies to commerce. Many agentic systems today try to automate shopping or travel planning end-to-end. A short-sighted approach optimises for cost and speed. But in doing so, we risk losing experience, loyalty, and trust - not to mention the strain that AI-driven micro-transactions could place on broader ecosystem efficiency.&lt;/p&gt;

&lt;p&gt;What if AI’s role were instead to curate, inspire, shortlist, or create a shopping list? This leaves the final say to humans where decisions and precision are required. That would take us much further, while laying a solid foundation for AI to operate responsibly and leaving room for new innovations to emerge.&lt;/p&gt;

&lt;p&gt;We should not try to pour new AI wine into old process casks. Let AI guide us toward new customer journeys, experiences, and ways of working.&lt;/p&gt;
</description>
        <pubDate>Fri, 20 Mar 2026 06:29:41 +0000</pubDate>
        <link>https://geord.ee/2026/03/new-wine-in-old-wineskins</link>
        <guid isPermaLink="true">https://geord.ee/2026/03/new-wine-in-old-wineskins</guid>
        
        
        <category>blog</category>
        
        <category>thoughts</category>
        
      </item>
    
  </channel>
</rss>
