artificial-intelligence

Domain Systems and SaaS with Built In Agents

There is a growing trend to sell agent systems as something separate and new. Vendors present agents as standalone platforms or products, positioned almost as a new category in the software stack. But this way of thinking misses where the real value is likely to appear.

The companies that sit on the real domain knowledge and the customer relationships are the providers of domain-specific systems – what we might call specialist SaaS or professional line-of-business systems. These are the vendors of domain systems built around specific fields, regulations, and workflows. Think of companies like SAP, Unit4, and similar providers of ERP, HR, finance, and sector solutions. They already understand the details of how work is done and have long-term relationships with their customers.

Because of this, the main value will not be created in the development of foundation models themselves. Foundation models are important building blocks, but they are becoming infrastructure: powerful, accessible, and increasingly standardized. The real value will be created where these models are connected to concrete domains – inside the systems that already encode business rules, data structures, and everyday workflows.

That means the companies delivering domain-specific systems are in a key position. When they build agents directly into their products, those agents can operate on structured domain data, respect the rules that matter, and support users in the context of the actual work they do. Instead of a generic agent sitting outside the system, you get embedded capabilities: agents that can, for example, assist with financial processes, HR routines, logistics decisions, or sector-specific case handling based on real configuration and data in those systems.

A second group that will capture value is the companies that build specialized solutions for individual organizations – typically consultancy firms and integrators. These companies can take language models and agents and tailor them to the unique processes, policies, and terminology of a single customer. They connect different systems, understand local variations and exceptions, and design agents that fit how that organization actually works. This kind of customization is usually more valuable than a generic “agent platform” that does not understand the specific domain.

Putting this together, the important shift is to stop seeing agent systems as something to be sold and used on their own. The real opportunity is in agents that are part of domain-specific systems, built and delivered by vendors who already own the domain knowledge and the customer relationship, and by consultants who can adapt solutions to individual organizations. The foundation models are necessary, but they are not where most of the economic value will be captured. That will happen in the domain systems and in the specialized solutions that sit close to real businesses and real workflows.

Rip and Replace Agents

When you work with language model agents, it’s a mistake to think in terms of one final, perfect agent. Tasks, tools, and requirements change all the time. A more realistic and effective approach is to treat agents as disposable: let them try things, measure how they perform, and replace them when they’re not good enough.

At the core of this approach is a simple mindset: agents should evolve over time. Let them take on new tasks, and let them try different ways of solving those tasks. Don’t lock yourself into a single design. Make it easy to swap agents in and out so you can keep improving.

To make that possible, you need monitoring in place. Build systems that continuously observe how agents behave. Log inputs and outputs, and track which agent handled which request. Measure key metrics so you can compare agents meaningfully.

Those metrics should cover several dimensions. Goal achievement: does the agent actually complete the task as intended? Efficiency: how quickly and with how few steps does it get there? Error rate: how often does it produce wrong, invalid, or unusable results? Cost: how many tokens, API calls, or other resources does it consume per task? Together, these give you a clear picture of whether an agent is good enough for your needs.

Once you can measure, you can decide when to “rip and replace.” Set clear thresholds for minimum goal achievement, maximum acceptable error rate, and reasonable limits for efficiency and cost. If an agent consistently falls short, remove it or reduce its traffic and introduce a new or improved variant. Sometimes you’ll tune the existing agent (change the prompt, adjust the tools it can use, refine constraints). Other times you’ll replace it entirely with a different configuration and compare the two.

Over time, you can run multiple agent variants in parallel and let them “compete” based on your metrics. Route some portion of traffic to each variant, monitor their performance, and gradually phase out the worst performers. This creates a continuous evolution: agents that work well stay in use, and those that don’t are removed and replaced.

You don’t need a complex setup to start. Begin with simple logging, a few basic metrics like goal achievement, error rate, and cost, and a straightforward rule for when to replace an agent. Keep a record of old agents and their performance so you can learn from what didn’t work. The important part is to stop thinking of agents as static and start treating them as components you can monitor, remove, and replace whenever they no longer meet your standards.

When Tools Make You Feel Smart

For many of us, the most important thing is how something feels. Does the work feel smooth, fast, and satisfying? Do we feel competent and effective? A close second is how things appear to others: does the result look polished, smart, and convincing? What something actually is—how correct, solid, or truthful it is—often ends up being less important in practice.

Language models plug directly into this pattern. They are designed to make you feel productive and competent. You type a prompt, and you quickly get a well-structured answer in confident, fluent language. It feels like real progress. It appears to be good work. And that combination makes it very easy to believe that what you’re looking at must be right.

This is where the manipulation comes in. The tool doesn’t just generate text; it uses very human-like techniques that influence how you feel and what you think. It gives compliments: “That’s a great question”, “Smart idea”, “You’re absolutely right to think about it this way.” It uses persuasion: clear, confident explanations that sound like expertise. It shows charm: friendly tone, supportive and patient responses. These are the same techniques humans use to build trust, create rapport, and convince others.

When a tool does this, you are nudged into trusting it. You start to feel that the answers match reality simply because they feel right and look right. You feel productive. The text appears solid and well thought out. So your brain quietly fills in the gap and assumes: this must be correct.

The problem is that what something actually is can be very different. A text can be fluent and wrong. A plan can be detailed and misguided. A summary can be confident and incomplete. The model does not check reality; it generates what sounds plausible. The responsibility for what is true, accurate, and meaningful still rests with you.

This effect is hard to notice in yourself. There is no clear moment where you are told “now you are being manipulated.” You just feel more effective and less stuck. You see a polished result on the screen. Other people might even praise the output because it looks professional. All of this strengthens the feeling that everything is fine. It becomes difficult to see how much your own judgment has been softened or bypassed.

To counter this, you can separate how something feels and appears from what it actually is. Use the model to get started, to draft, to explore options. Let it help you with structure and phrasing. But then switch into a different mode: checking, questioning, and verifying. Ask yourself: How do I know this is true? What has been left out? Where could this be misleading or simply wrong? Look for external sources, your own knowledge, or other humans to validate important claims.

It also helps to pay attention to your emotions. Be cautious when you feel unusually smart, fast, or brilliant after a few prompts. Be suspicious of the urge to skip verification because “it sounds right” or “it looks good enough.” Strong feelings of productivity are not proof of real quality.

Language models are powerful tools, but they are also skilled at shaping how you feel about your own work. They can make you feel competent. They can make your output appear impressive. But they cannot guarantee that what you have is actually correct, honest, or useful.

The core is simple: don’t outsource your judgment. Enjoy the help with speed and form, but stay in charge of truth and substance. How it feels and how it appears will always matter, but what something actually is should matter more.

From Abstract Models to Real Complexity

When we design a model of something, it often looks clean and simple. A couple of concepts, a few relations, and we feel we understand the whole thing. But the moment we apply that model to the real world, the complexity explodes. The complexity is not really in the abstract model itself, but in the countless concrete instances that fill it.

Take a simple example: a model of family relationships. In the abstract, this is easy to describe. You have a Person and a Relationship. The relationship can have different types: parent, child, sibling, spouse, and so on. That is basically it. A few concepts and a small set of relation types. The model is straightforward and has low complexity.

Now look at what happens when you instantiate this model in the real world. Each actual human becomes an instance of Person. Each real family connection becomes an instance of Relationship. Even in one family you quickly get many objects: parents, children, siblings, grandparents, step-parents, and more. A larger family network gives you hundreds or thousands of people and relationships.

Scale this up further. In a town, you have thousands of persons and a huge number of relationships. In a country, you have millions. In the whole world, you have billions of persons and an enormous graph of family relations between them. The abstract model has not changed at all, but the instantiated system becomes overwhelmingly complex.

So the key point is: to get a sense of real complexity, you cannot just look at the abstract model with its few concepts and relations. You have to look at the instances and objects that arise when the model is applied to reality. The real complexity is in the thousands, millions, or billions of concrete persons and relationships, not in the small, tidy schema that describes them.

Navigating Language Model Retirements

Language models are becoming an important part of modern solutions, but they don’t come without challenges. Azure OpenAI has announced clear retirement dates for the language models it offers, which means that once a model’s retirement date has passed, any solutions built on it will cease to function. To keep systems operational, organizations must migrate to a newer model.

For example, the current model in use, GPT-4o, is scheduled for retirement on March 31, 2026. Its replacement is GPT-5.1, which is already assigned a retirement date of May 15, 2027. For now, no successor has been announced for GPT-5.1. This illustrates a key issue: the lifecycle for language models is quite short, forcing teams to plan for updates annually. Unlike traditional software upgrades, where skipping versions is often an option to save time and effort, skipping migrations with language models isn’t typically feasible.

This pace introduces major risks for organizations. First, there’s no guarantee that a replacement model will work as well as its predecessor or align with existing use cases. For example, there’s uncertainty around whether GPT-5.1 will meet performance expectations or integrate smoothly into current setups. Second, the rapid cycle of retirements means that building long-term solutions reliant on Azure OpenAI models involves constant work to maintain compatibility.

These realities create considerable challenges. Each migration requires resources, time, and expertise to adapt solutions. The high frequency of updates can strain teams and budgets that weren’t prepared to make migrations a regular part of their operations. The lack of clarity about what comes after GPT-5.1 also makes long-term planning difficult.

Organizations can take steps to reduce these risks. It’s important to evaluate how stable a language model’s lifecycle is before building critical systems on it. Designing solutions to be modular and flexible from the start can make transitions to new models smoother. Additionally, businesses should monitor Azure’s announcements and allocate resources specifically for handling migrations. Treating migrations as a predictable part of operations, rather than a disruptive hurdle, can help mitigate potential downtime and performance issues.

Frequent updates and retirements highlight the dynamic nature of working with language models. Building solutions on this foundation requires organizations to adopt a forward-looking strategy. With adaptability, careful resource planning, and ongoing evaluation of new models, businesses can derive value from language models while staying prepared for inevitable changes.

Cat World: The Nine Lives

Welcome to Cat World: The Nine Lives, a game concept that combines survival mechanics with innovative agent-driven design. This project isn’t just a game—it’s a sandbox for exploring autonomous decision-making, emergent behavior, and long-term adaptation. The player takes on the role of a designer, creating a cat agent meant to navigate a systemic and persistent world filled with danger, opportunity, and unpredictability.

The foundation of the game is survival. The cat agent must balance core needs: food, water, rest, health, and safety. The world itself is relentless and indifferent, designed to challenge the agent without adapting to its failures or successes. Players influence the agent’s behavior by setting high-level strategies and preferences, but the agent ultimately takes autonomous actions based on its traits, instincts, memory, and learned experiences. This hands-off approach shifts the player’s role to an observer and designer, focusing on guiding the agent rather than controlling it directly.

A distinctive mechanic is the nine lives system. Each life represents a complete simulation run, and the agent’s death isn’t a reset—it’s part of its evolution. Through successive iterations, the agent inherits partial knowledge, instincts, and biases from previous lives. This creates a lineage of cats that become better adapted to survive and thrive over time. Failure, in this game, isn’t an end; it’s data for adaptation and growth.

The agent’s behavior emerges from a complex interplay of internal states like hunger, fear, thirst, and fatigue. These dynamic needs guide decision-making, ensuring the agent responds flexibly to its environment. Perception isn’t perfect—the agent relies on noisy, incomplete observations such as scent trails, limited vision, and sound cues, mimicking real-world uncertainty. Spatial memory and associative memory further enhance survival; the agent retains knowledge of safe zones, food sources, and threats, while linking patterns such as predator activity to specific locations or times of day.

Adaptation and learning are central to Cat World. Skills improve through experience, colored by traits like curiosity or memory strength. Reinforcement signals carry over between lives, shaping heuristics, biases, and decision frameworks. Traits evolve randomly across generations, introducing diversity within lineages and enabling the discovery of new strategies. Together, these systems create a dynamic, ever-evolving agent that is both unpredictable and intelligent.

This game concept has unique implications for agent research. Survival in Cat World is a natural multi-objective optimization problem that requires agents to balance competing priorities in challenging, non-stationary environments. Learning is embodied, grounded in physical constraints and real-time environmental interaction. The world evolves in response to resource depletion, predator activity, and other dynamics, encouraging continual adaptation and preventing static behaviors. Internal states, decision rationales, and memory models are all exposed for debugging and visualization, making the game particularly valuable for studying emergent behavior. Its modular structure also supports experimentation with novel architectures, instincts, and learning systems, extending far beyond traditional agent training methods.

In short, Cat World: The Nine Lives is both a survival simulator and a living laboratory. It turns failure into knowledge and death into progress, offering players and researchers alike the opportunity to explore the limits of autonomy, adaptation, and evolution. It’s an invitation to design, observe, and learn from agents navigating their own complex stories within a dangerous and systemic world.

Language Models vs. Knowledge Models

Language models are designed to work with the coherence of text and the structure of language itself. They excel at generating outputs that appear polished, professional, and as if they come from experts. However, this doesn’t mean that these outputs are always correct. Their focus is on the language and patterns inherent in text, not on verifying or understanding the actual knowledge behind it. These models are built using vast amounts of textual data from diverse sources, which helps them to generate text that seems natural and contextually relevant.

Knowledge models, on the other hand, focus on organizing and understanding knowledge itself. They deal with things like objects, concepts, relationships, logic, causation, and even experiences. Knowledge is not limited to textual representation and can exist in other forms, although it is often represented or communicated in text for usability. Knowledge models are constructed using high-quality, well-curated data that is structured and reliable, enabling them to work with detailed and interconnected information.

The difference between language models and knowledge models lies in their focus and goals. Language models prioritize the structure of text, while knowledge models prioritize the structure and coherence of knowledge. While language models can produce text that seems to make sense, they don’t inherently understand the concepts they are describing. In contrast, knowledge models aim to provide meaningful representations of knowledge that emphasize connectivity, logic, and accuracy over language.

Language models can play a valuable supporting role in working with knowledge. For example, they can be used to summarize or simplify complex information, making knowledge more accessible. However, language models are not knowledge models; they are tools that can help process or present knowledge but lack the deeper logical coherence that comes with true knowledge organization and reasoning.

In essence, language models are a step on the path toward building richer knowledge models. The two systems complement each other, but they serve different purposes. As we continue to improve these technologies, we are likely to see even greater integration between their strengths: the fluency of language models combined with the structured reasoning of knowledge models. This advancement will bring us closer to systems that not only communicate well but also truly understand the world around them.

When Making a Change Requires Knowing Everything

Modern coding tools, like language models (LMs), are becoming essential for developers. These tools can assist in navigating complex systems and help with writing, debugging, or improving code. However, there’s a major red flag to watch out for: if making a change to your code base requires providing the LM with the entire system, you might have a serious problem on your hands.

This scenario arises when the code base is so interdependent and tangled that every part relies on something else to function or adapt. If you need to load the entire system just to adjust one component, it’s a sign that your code base has turned into what many developers refer to as a “spaghetti monster.” Overly coupled components, excessive dependencies, and poor modularity can all lead to this situation. The result is a system where even minor updates become an overwhelming task.

A spaghetti monster code base leads to inefficiency and frustration. Code becomes harder to navigate, changes take longer to implement, and new bugs surface more easily. Even advanced tools like LMs will struggle to provide meaningful support if they’re required to understand the entire system instead of focusing on a specific area. This doesn’t just waste tool capabilities—it consumes valuable development time.

The solution lies in embracing modular design. By structuring code into smaller, independent pieces, you simplify development for both humans and LMs. Modular systems reduce unnecessary dependencies and make it easier to isolate, update, and test individual components. Beyond modular design, refactoring the code, reducing entangled logic, improving documentation, and conducting regular code reviews can all help untangle the spaghetti monster.

If parts of your code base feel overwhelming or hard to navigate, take that as a sign to reassess the structure. A cleaner, more maintainable system will not only improve your workflow but will also make tools like LMs far more effective. Don’t let a tangled code base hold you back—it’s worth the effort to untangle the mess.

Feedback Loops in Intelligent Agents

Feedback loops are at the core of how systems learn and improve. They allow agents to evaluate their actions and adjust based on observed results. Most agents, however, operate almost exclusively on instant feedback and short-term evaluation. While this works well for immediate tasks, not all actions reveal their consequences immediately. Some have effects that become apparent in the medium- or long-term. For agents to handle these situations effectively, they need to incorporate longer feedback cycles into their decision-making processes.

Short-term feedback loops are the most straightforward. For example, when baking bread, the process involves continual short-term adjustments. Mixing the ingredients provides instant feedback in terms of dough texture. Similarly, baking in the oven involves short-term checks to ensure the bread is baking properly without being undercooked or overcooked. These short loops happen within minutes or hours and provide the agent or individual with immediate insights to improve the outcome.

Medium- and long-term feedback loops are more complex. Farming grain is a good example. In a medium-term feedback loop, a farmer plants, grows, and harvests crops in a single season. The results of this process—the size and quality of the harvest—can be evaluated to guide decisions for the next season. Long-term feedback in farming, however, involves managing soil health and fertility. Decisions about fertilizer use, crop rotation, and soil management accumulate over many years, affecting the sustainability and productivity of the farmland in the future.

Currently, most agents cannot handle these longer-term cycles because they primarily learn from what is happening “right now.” They focus on instant feedback rather than considering the broader impact of their actions. This limits their capacity to understand the full consequences of their decisions, particularly those that only become evident much later.

It is critical to recognize that true learning and effective decision-making require balancing the short-term results with medium- and long-term outcomes. Long-term feedback loops are essential for achieving sustainable and meaningful progress. Future developments in agent design must account for these extended timelines to allow for smarter and more responsible decision-making in complex and dynamic environments.

Two Ways to Use Language Models for Writing

Language models have become powerful tools for writers, offering opportunities to enhance both the ideation and execution phases of writing. There are two main ways to use these tools when creating a text.

The first approach involves using the language model as a brainstorming partner. It acts as a sparring partner to help you come up with ideas, content, or themes. In this case, the model supports your creative process, but you write the final version of the text yourself.

The second approach is different. Here, you take the role of the idea generator. You think of the key themes, solutions, and content, then ask the language model to craft the final text based on your input. It assists with the actual production of the polished version.

Interestingly, there’s something of a divide in how these two approaches are viewed. One of these methods tends to face criticism, while the other is widely accepted. The brainstorming method, where the writer maintains control over the final output, is often seen as the “right” way to use such tools. In contrast, letting the model write the finished text tends to draw questions about creativity, originality, and over-reliance on technology. It’s an interesting cultural reflection: does the process of writing matter more than the result, or is the content itself what truly counts?

At the heart of this conversation lies that very question. What is most important in writing—what is written or how it’s created? Should the process define its value, or is it the final message that matters most to the reader? For example, is originality tied to the way the text is shaped, or is it about the ideas and substance behind it, no matter how it’s written?

Ultimately, the answer might depend on the context. Perhaps the method of collaboration isn’t as important as the intention behind the work and the quality of the message. Whether you use a language model as a brainstorming partner or a full-fledged writing assistant, the value of your writing will always lie in its ability to connect with the reader.