Microsoft Agent Framework: What I Stopped Building Myself

The LLM call was the easy part. The execution loop around tools, context, retries, approvals, and state was not.

I've built several AI workflows where the LLM call itself was the easy part. The code around it was not. Once an agent needs tools, state, context, retries, approval points, and some control over what happens next, a simple model call starts turning into a small runtime. I ran into this while working on an AI inspection workflow for Moonraker. The system handles parts of a property inspection such as plumbing, waterproofing, tiles, and walls. The first version used orchestration code I controlled myself. It worked. That was not the problem.

The problem was that more and more of my code had nothing to do with property inspections. I was building agent infrastructure.

That was when Microsoft Agent Framework started to make sense for me.

The agent wasn't the complicated part

The original flow was straightforward at first. An inspection request entered the application, I prepared the context, called the model, processed any tool calls, updated the state, and decided what should happen next.

Then the normal production questions started appearing.

What if a tool fails? What state should the next model call receive? How do I stop a bad execution loop? Which actions should require approval? How much conversation history should I keep? How do I understand what happened when the final result is wrong?

Each answer adds another small piece of infrastructure.

Before the migration, the application owned both the inspection logic and most of the generic agent execution plumbing

Figure 1. Before the migration, the application owned both the inspection logic and most of the generic agent execution plumbing.

None of these pieces is especially difficult on its own.

The problem is that they accumulate.

A tool call needs validation. Then it needs failure handling. The result has to become part of the next context. Some actions should not execute automatically. The workflow needs to know when it is finished.

Eventually, code that only keeps the agent running starts to crowd out the interesting part of the application.

I was slowly building my own agent harness

Looking back at the implementation, that is basically what was happening.

I had application code responsible for the execution loop, tool invocation, context management, state transitions, failure paths and the rules around continuing or stopping execution.

A simplified version looked conceptually like this:

while not finished:
    context = build_context(state)

    response = await model.invoke(context)

    if response.has_tool_call:
        result = await execute_tool(response.tool_call)
        state.add(result)

    elif requires_approval(response):
        decision = await request_approval(response)
        state.add(decision)

    else:
        finished = should_finish(response, state)

Real code becomes less clean very quickly.

Retries appear. Tool errors appear. State gets larger. You add tracing because the final answer is wrong and the cause happened several model calls earlier. Then another special case enters the loop.

There is nothing fundamentally wrong with implementing this yourself.

For a small AI feature, I still prefer explicit application code over adding a framework too early.

Eventually I had to ask:

Is this orchestration part of my product, or am I rebuilding generic agent infrastructure?

That distinction changed the architecture.

What Microsoft Agent Framework actually replaced

Microsoft describes an agent harness as the scaffolding around the model: an execution loop with tool invocation, planning, memory, context management, approvals, and telemetry.

That matched the category of code I increasingly did not want to maintain myself.

The migration was not about replacing prompts or handing the whole application to a framework. It was about moving generic agent mechanics behind a reusable runtime while keeping the inspection-specific behavior in the application.

After the migration, the framework owns more of the generic execution machinery while the application keeps the domain decisions

Figure 2. After the migration, the framework owns more of the generic execution machinery while the application keeps the domain decisions.

It sounds like a small architectural change, but it changed where I spent my time.

Instead of adding another branch to a custom execution loop, I could spend more time on what the inspection system was actually supposed to do.

What I stopped building myself

The main win was not a new AI capability. I could already build the workflow.

The win was removing responsibility for pieces of infrastructure that were not specific to Moonraker.

Microsoft Agent Framework's harness provides a runtime around a chat client. It includes function invocation, per-service-call history persistence, context compaction, planning support, memory, tool approvals, and OpenTelemetry-based telemetry. I can configure those capabilities individually instead of accepting one fixed setup.

That matters because these are exactly the kinds of concerns that start appearing as an agent becomes more than a prompt with tools.

I no longer want to write another generic tool-calling loop just because I started a new agent.

I also don't want every project to invent its own answer to context growth, execution history, approval plumbing and telemetry.

Those are useful capabilities. They are just not where my product is differentiated.

What I did not give to the framework

The more important question is which orchestration I still own.

I still need to decide:

  • what an inspection agent is allowed to do
  • which tools exist and what each tool can access
  • what belongs in application state
  • which operations require deterministic validation
  • where model reasoning is useful
  • where a normal if statement is better
  • what a valid inspection result looks like
  • which actions need human approval
  • what happens when a domain rule conflicts with an agent decision

Those are product and domain decisions.

I don't want a framework making them for me.

| Microsoft Agent Framework | Moonraker Application | | ------------------------- | ------------------------- | | Agent lifecycle | Business rules | | Tool invocation mechanics | Permissions | | Context handling | Inspection state | | Approval infrastructure | Tool implementations | | Telemetry hooks | Validation | | Common agent plumbing | Deterministic constraints | | | Product behavior |

Figure 3. The useful boundary for me: framework infrastructure on one side, domain behavior on the other.

The boundary is not absolute, but it is more useful than saying "the framework handles the agent."

The application still handles the agent where it matters.

Less code was not the goal

It is tempting to measure a migration like this by lines of code.

I don't think that tells me much.

I can replace 200 understandable lines with one library call and still make the system harder to debug.

The question I care about more is:

Which complexity do I actually want to own?

For domain logic, I want explicit control.

If an inspection result must pass a specific validation rule, I would rather express that rule in application code than hope the model consistently infers it from a prompt.

If a tool must never run without permission, that should be enforced as a boundary.

If an output must match a schema, I validate it.

If the next workflow transition is already known, I don't need an LLM to rediscover it.

But maintaining another implementation of a generic agent loop gives the product very little advantage.

That is complexity I am comfortable moving into a framework.

The tradeoff: my code disappeared, the complexity didn't

There is an obvious downside.

When I own the loop, I can follow every branch because I wrote every branch.

When I move that responsibility into a framework, part of the execution happens behind an abstraction.

The complexity has not vanished. It has moved.

That makes observability essential.

If an agent chooses the wrong tool, repeats a step, builds the wrong context or produces an unexpectedly expensive execution path, I still need enough information to reconstruct what happened.

Microsoft Agent Framework has OpenTelemetry support for agent activity, and the harness emits telemetry around its execution. That is useful, but I still need to decide where those traces go, what I monitor, and what constitutes abnormal behavior in my application.

The framework can expose the execution.

It cannot decide what good execution means for my product.

Agents still need deterministic boundaries

This migration reinforced something I have seen in other AI systems: I don't want the model deciding everything.

LLMs are useful where the problem contains ambiguity.

They can interpret an inspection description, reason over unstructured information, decide which relevant capability to use, and synthesize results.

But if the system already knows a rule, I prefer code.

The agent operates inside deterministic application boundaries rather than replacing them.

Figure 4. The agent operates inside deterministic application boundaries rather than replacing them.

For production AI, I want this shape:

application → controlled agent runtime → model and tools → validation → application

The agent operates inside the application.

The application should not operate inside the agent.

Human-in-the-loop is architecture, not a prompt

Approval is a good example of why this separation matters.

There is a big difference between telling a model "ask the user before doing X" and designing the workflow so execution actually pauses until an external approval arrives.

Microsoft Agent Framework workflows support request/response handling for this kind of human-in-the-loop interaction. An executor can issue a request to an external system and wait for a response before the workflow continues.

I want sensitive actions controlled at that level: as an execution boundary, not polite text in a system prompt.

This becomes increasingly important as agents move from generating answers to performing actions.

Context management becomes infrastructure surprisingly fast

Context is another piece I initially treated as application logic.

For short workflows, manually controlling history is easy.

Longer tool-calling sessions are different. Tool outputs accumulate. Previous model turns accumulate. Instructions compete with execution history. Eventually the context itself becomes something the runtime has to manage.

The current Agent Framework harness includes context compaction for long-running sessions. Good context engineering is still necessary, but it separates two problems that are easy to mix together:

  1. What information does my agent need?
  2. How does the runtime keep a growing execution within the model's context budget?

I still own the first problem.

I am happy for infrastructure to help with the second.

That is another example of the same boundary.

When I would use Microsoft Agent Framework again

I would not introduce an agent framework just because an application calls an LLM.

If the system is basically:

I don't need one.

Even a small number of deterministic tool calls can often be easier to understand as normal application code.

I start considering an agent framework when several of these things appear together:

  • repeated model/tool interactions
  • execution state across multiple steps
  • multiple tools or agents
  • human approval points
  • branching workflows
  • long-running execution
  • context that needs active management
  • recovery from interrupted execution
  • enough orchestration code that it starts looking like its own subsystem
  • a real need for traces across model and tool activity

At that point, the abstraction starts paying for itself.

Before that point, it can easily become another dependency you have to understand.

A framework can remove complexity or become the complexity. Recognizing that line matters more than choosing the framework with the longest feature list.

The part I would still build myself

If I started the Moonraker inspection workflow again, I would use the framework earlier for the generic agent runtime.

I would still keep the same skepticism about putting business behavior inside it.

The framework can manage the mechanics of an agent calling a waterproofing inspection tool.

My application should decide whether that tool is available for the current inspection, what data it can receive, how its result is validated, whether the result changes the inspection state, and whether a human needs to approve the next action.

Those decisions make the system my application, not a generic agent.

And that is the code I want to own.

What actually changed for me

Moving to Microsoft Agent Framework did not suddenly let me build something that was impossible before.

I already had a working workflow.

What changed was the amount of generic agent infrastructure I was willing to keep maintaining myself.

It is a less dramatic reason to adopt an agent framework, but it is the useful one for me.

The LLM call was never the difficult part.

The difficult part was everything required to make an LLM behave like a controlled component inside a larger software system.

I still own that system.

I just don't think I need to own every piece of plumbing underneath it.

References

  • Microsoft Agent Framework: The Agent Harness
  • Microsoft Agent Framework at BUILD 2026
  • Microsoft Agent Framework Workflows: Human-in-the-loop
  • Agent Harness: Making your agent production-ready