
Agent Interoperability and Tool Protocols: What Actually Needs to Be Standardized
A protocol standardizes the shape of a call. It does not standardize what the call means, who is allowed to make it, or what happens when it fails.
Read reportSystems that use models to plan or execute multi-step tasks through tools, memory, or other agents, including their design, evaluation, and deployment.
Tagged articles
30 articles in this tag.

A protocol standardizes the shape of a call. It does not standardize what the call means, who is allowed to make it, or what happens when it fails.
Read report
An agent with legitimate access to a folder can still write four thousand files somewhere it should not, and every access check will pass.
Read report
That is the uncomfortable lesson from AutoJack, a chain Microsoft's security team disclosed in June 2026. The individual bugs were ordinary — the kind of…
Read report
The agent that answered cleanly for ten minutes starts contradicting a decision it made at turn twelve. It asks again for a fact you supplied in the…
Read report
A failed response is not a failed operation. A successful response is not a successful operation either. Everything hard about AI agent reliability lives…
Read report
An audit trail is a reconstruction contract. If a reviewer cannot explain a consequential action from the record alone, you have activity logs, not…
Read report
The attacker's cost curve is falling faster than the defender's learning curve. That single asymmetry explains most of what is actually changing.
Read report
The demo passed. The patch was clean. Three weeks later, your team spends more time reviewing agent output than it would have spent writing the code by…
Read report
The diff counter climbs every sprint. The release cadence does not move. That gap is not a tooling problem — it is a measurement problem, and most teams…
Read report
The unit of work is shifting. Autocomplete suggested a line; a coding agent takes a goal, touches the file system, runs a command, reads the error, and…
Read report
A deflected contact is a closed conversation. A resolved issue is a problem that stayed solved.
Read report
Your pager fires at 2:14 a.m. A customer has posted a screenshot: your assistant told someone to adjust a medication dose. You open the trace, find the…
Read report
The hard part was never the model call. It is everything you have to build around it.
Read report
Your application code can be perfect and your AI system can still be compromised before it ever runs. The model weights, the fine-tuning dataset, the MCP…
Read report
A five-step task at 85% per-step accuracy finishes correctly about 44% of the time. That arithmetic, not the model, is what decides whether your workflow…
Read report
The cluster is provisioned, the dashboards are green, and the inference bill is still climbing while latency drifts. That combination — healthy…
Read report
Writing code stopped being the bottleneck. Reviewing and integrating it became one.
Read report
The prompt was precise. The model was capable. The change still came back wrong.
Read report
A coding agent is not a smarter autocomplete. It is a shell with your credentials, your network, and your filesystem — and it runs all three before you…
Read report
The agent returns a 900-line diff across eleven files. Tests pass locally. The reviewer opens the pull request, scrolls twice, and still cannot say whether…
Read report
A computer-use agent is a model that reads a screen and drives a mouse and keyboard. The demo looks like magic. The second run looks like a different…
Read report
The demo proves what a model can do. Deployment proves what a system will let it do. An autonomous agent can plan, call tools, and act toward a goal on its…
Read report
Agentic AI systems—autonomous agents that plan, call tools, and adapt as they work—are moving out of research demos and into business-critical workflows.…
Read report
A benchmark score is a result under agreed test conditions. Reliability is what remains when ordinary inputs, missing data, delays, and recoverable failure…
Read report
You built an agent. It worked once. Then you changed the input slightly and it fell apart, and you could not tell which part of the system failed.
Read report
The real question is not whether to use many agents. It is what workflow constraint justifies adding another agent, and what coordination contract makes…
Read report
An AI agent is software that uses an AI model to choose actions, call tools, and work toward a goal. Agentic AI describes systems built from these agents,…
Read report
The attacker never talks to your model. They leave a sentence in a document it will read later.
Read report
A voice agent that transcribes every word correctly can still feel broken. The transcript is not the conversation.
Read report
A working demo and a working workflow are different artifacts. One proves the model can do the task. The other proves the organization can repeat it.
Read report