Ten years ago, the companies that moved to cloud-native compounded for a decade. The ones that did not are still paying for that decision. The same window has just opened for AI-native, and it opened in the last twelve months.
For most of the last three years, generative AI inside companies has been a moving target. You would commission an internal tool, your team would build it on whatever framework was hot that quarter, and six months later, half of it would be obsolete. That is not a useful environment for serious investment, and most boards correctly treated it as exploratory spend.
That has changed. Four things have settled into place at roughly the same time: a stable set of open standards for how agents work, an infrastructure layer for agents to discover and act on, a set of engineering practices for keeping probabilistic systems reliable, and interfaces that connect the whole thing to humans and to other organizations. Each has been emerging for years. They have only just landed together.
I would put this on the same order of significance as the shift from on-premises to cloud. Cloud-native was a bet that, if you built your software around containers, managed services, and elastic infrastructure, it would compound. AI-native is the equivalent bet around agents, tools, and self-improving workflows. The companies that make the transition early will spend the next decade compounding on it. The ones that do not will find themselves competing against organizations that have rebuilt their cost structure underneath them.
This post is a map of what AI-native actually means in practice, not as a marketing label, but as four concrete investments. I have been part of a team that built most of these in production. The newest pieces, the agent mesh in particular, we are working on right now. Where I am reporting from experience, I will say so. Where I am reporting from the frontier, the same.
The four investments
Treat the rest of this post as a single argument with four parts.
Standards are the open formats and protocols that finally make agentic systems portable across vendors and across time. They are the reason your investment now has a chance of still being useful in five years.
Infrastructure is the substrate on which those standards sit: self-describing APIs, a data foundation, and an emerging agent mesh. Without it, your agents have nothing to discover or act on.
Practices are the engineering disciplines that make probabilistic systems behave reliably enough to bet on. They are how you get from “the demo worked” to “we can put this in front of customers.”
Interfaces are what make the whole investment legible to the rest of the business: a cockpit where your people supervise swarms of agents, and an outward-facing endpoint where your agents talk to your customers’ agents.
None of these is optional. Standards without infrastructure give you agents with nothing to do. Infrastructure without practices gives you a probabilistic system you cannot trust in production. Practices without interfaces give you a system nobody uses. The argument for “why now” is that, for the first time, all four are simultaneously buildable.
1. Standards
A protocol is only useful if multiple vendors agree on it. Three have, in the last year.
MCP (Model Context Protocol) is how agents call tools. Introduced by Anthropic in November 2024, it now has well over 16,000 servers in the wild, with OpenAI, Google, GitHub, Linear, Replit, Zapier, and most other major vendors integrated. You can build an MCP server for your internal system today and reasonably expect it to be useful to any agent your team picks up next year.
A2A (Agent-to-Agent) is how agents talk to each other. Announced by Google in April 2025, donated to the Linux Foundation in June 2025, and now supported by more than 150 organizations, including Microsoft, AWS, Salesforce, SAP, ServiceNow, and IBM. This is the standard that makes cross-organization agent communication tractable.
Agent Skills is how you package expertise so an agent can use it. A skill is a folder containing human-readable files that a domain expert can write, version, and review without engineering help. Introduced by Anthropic in late 2025, released as an open standard, and adopted by Codex CLI, Gemini CLI, Cursor, and several others.
You will notice these three do not overlap. One for tools, one for agents, one for knowledge. Each has cross-vendor adoption. Each is supported by an open standards body. Each has survived the year-long obsolescence cycle that killed previous generative AI architectures. That is the actual argument for the word “standard”: these are the pieces that will still be there next year.
The practical implication is straightforward. You can now tell a competent team, “Build this using MCP, A2A, and Agent Skills,” and the odds that what they ship is still valuable in twelve months are dramatically higher than at any previous point in this cycle. That alone is a regime change.
2. Infrastructure
Standards are building blocks. They need something to be built on.
Self-describing APIs are the prerequisite nobody talks about. The discipline of making your APIs human and machine-readable from the description has been a good practice since Open API has established itself as the way to define RESTful APIs. MCP doubles down on it. The quality of your tool descriptions directly determines whether an agent can use your system. If your APIs are not already documented to that standard, that is the first piece of debt to clear. Your documentation is now a capability, not a hygiene item.
A modern data foundation is the unglamorous part of the substrate, and the one most often skipped. Whether you call it a data mesh, a data platform, or something else, the principle that matters for agentic systems is federated governance. Agents have to be able to check at runtime whether they are authorized to read a given data product, without a central team having to prewire every permission. Combined with a catalog rich enough to be queried, you get the property that makes agents useful at scale: they can discover the data they need, verify they are allowed to use it, and publish new data products back.
If your organization has already invested here, you have a head start. If you have not, the work runs in parallel with the agentic build. It does not block you.
The agent mesh is the youngest layer and the one currently materializing across vendors. Solo.io, Solace, Lyzr, Databricks, and Microsoft have all proposed flavors of it. The clearest way to think about it is the cloud-native moment repeating for agents. What Kubernetes did for containerized microservices (scheduling, identity, policy enforcement, observability), the agent mesh does for agents. An agent gateway centralizes traffic between agents, models, and tools, enforcing authentication and audit. An agent registry catalogs which agents exist, what they can do, and who owns them. An agent runtime schedules them and handles the lifecycle.
This piece is the newest of everything in this post. My team and I are working on the transition. The shape is clear enough to belong on the map, with the honest caveat that the vendor landscape will evolve over the next year.
3. Practices
Agentic systems are probabilistic. The same prompt can produce meaningfully different answers, and you cannot eliminate that. You design around it.
Each practice below is a strategy for getting reliable behavior from a system whose components are individually unreliable. Without these practices, what you have is a demo. With them, what you have is production.
Process feedback is the highest-leverage practice and the one most often missed. The misconception goes like this: “AI systems learn from interactions.” They do not, not on their own. The model is retrained on a schedule set by the vendor. What you control is the data you collect from those interactions, and process feedback is the discipline of collecting it deliberately.
Every time your AI system makes a decision that a human reviews, a triaged email reassigned to a different queue, a priority ranking the user reorders, a suggested resolution the expert overrides, that human action is a labeled training example. Free. And especially valuable, because it is a case the system got wrong. Design your interfaces so the correction is easy for the human to apply and is recorded structurally for you. We have used this pattern in production, walking backward through existing workflows detailed enough to extract triaging signals from. It works.
Classical machine learning inside agentic systems is what closes the loop. Use the language model as the brain that decides which tool to call, and call a classical model as the tool that actually does the domain work. Those classical models are easy to fine-tune on the data you have been collecting via process feedback. Your company stops being a wrapper around someone else’s language model and starts being a real composition of domain expertise and general reasoning.
Human-in-the-loop gating is the complement to process feedback. For high-stakes decisions, you put a human approval step into the loop. Done well, this is not a bottleneck. It is a focusing device. The agent does the legwork, surfaces the decision, and your domain expert spends their attention on the part of the task that genuinely requires it. The byproduct is the same training signal as above.
Deterministic verification is the version of that gate where a program can check the answer, no human required. Compilers, type checkers, schema validators, anything that can decide “this output is correct” without ambiguity. Agents are mediocre at being right the first time and good at iterating. Deterministic verifiers give them a cheap signal to iterate against.
Self-actualizing memory is what enables agents to stop relying entirely on curated knowledge bases. After each interaction, the agent distills what was said into candidate facts, reconciles them with what it already knows (add, update, contradict, ignore), and pulls relevant memories back when needed. The reason this matters for the business is that your agent no longer needs a pre-curated knowledge store to be useful. It builds its own from the interactions it actually has. That changes the economics of deploying agents in places where there is no neat document corpus to point them at.
Agent-brokered asynchronous communication is, in my experience, the practice that unlocks the largest organizational gains. The pattern that consumes the most calendar time in traditional B2B work is synchronization between teams. Team A needs something from Team B, sends an email, waits, schedules a meeting, waits longer.
Replace the synchronous handoff with an agent. Team B’s agent has their domain knowledge encoded in skills, exposes their internal tools via MCP, and speaks A2A. Team A sends a request and gets an answer immediately for the majority of cases. The fraction that genuinely needs a human from Team B hits a human-in-the-loop gate, and a person from Team B is brought in for that decision specifically, not for the whole workflow.
The honest objection is that probabilistic systems can produce incorrect answers, and an asynchronous chain can act on them before anyone notices. That is exactly what the rest of this section is for. Gates on high-stakes calls, deterministic verifiers where possible, and process feedback, closing the loop. All of these exist precisely so async agent communication is safe enough to bet on. The payoff is not “remove humans from the loop.” It is: keep humans in the loop where their judgment is genuinely required, let asynchronous work flow at machine speed everywhere else, and keep the wisdom of different teams in the workflow without forcing everyone into the same meeting.
4. Interfaces
Standards, infrastructure, and practices are internal. They become valuable through two interfaces. One faces your own people. The other faces your customers.
Headless SaaS
This is the one with the most direct revenue implications, so it goes first.
The pattern: your B2B customers stop interacting with your product through its screens and its REST endpoints. Their agents talk directly to your agents. The classical interface recedes, and an A2A endpoint becomes the primary product surface. Customers ask your agent for things in natural language. Your agent does the work and reports back, with audit and policy in between.
This sounds speculative. It is not anymore. Industry analysts project that around 40% of enterprise applications will embed task-specific agents by the end of 2026. Teams on the agentic side already report that they interact with their own products through agents more than through the UI. If you are a B2B vendor and your customers’ agents cannot talk to your system, you will find them swapping you out for a vendor whose can.
This is the part of the AI-native bet with a competitive clock on it. The question is not whether your customers will eventually buy through agents. It is whether you are ready when they do.
The agent cockpit
Assume that no single agent does the whole job. Any meaningful workflow involves a swarm of agents collaborating, with humans intervening at the points the practices section identifies as high-stakes. The agent cockpit is the interface that lets a human supervise that swarm without becoming the bottleneck. The right design lets a domain expert zoom into the exact decision that needs them, make the call, and hand control back.
Done well, it is how your most expensive humans spend their time on the highest-leverage decisions and nothing else. Done poorly, it is yet another internal tool nobody opens. The difference is whether the cockpit treats human attention as a scarce resource to be earned, or as a default to be demanded.
What this means in practice
The argument is short enough to fit in a sentence: the technical pieces required to become an AI-native company are, for the first time, simultaneously stable enough to bet on.
That does not mean the work is easy. It means the work is now worth starting. A year ago, anything you built had high odds of being obsolete in months. Today, if your team builds on MCP, A2A, and Agent Skills, on a serious data foundation, with the engineering practices above, behind a cockpit and an A2A endpoint, there is a real chance that what they ship is still doing useful work in five years.
The shift to AI-native is on the same order as the shift to cloud-native a decade ago. The companies that moved early on cloud-native compounded for ten years on top of that decision. The companies that did not are still paying for it.
If you want a single concrete first move, here is one. Pick one workflow in your business where two teams currently synchronize through email and meetings, and where the cost of that synchronization is visible in cycle time. Audit whether the systems involved are agent-readable today. They probably are not. Closing that gap, for that one workflow, is small enough to do this quarter and revealing enough to tell you whether the rest of the journey is viable.
Start there.


