All articles
Architecture/October 5, 2026/7 min read

Building an Agentic System in .NET, Part 5: Designing an Agent Bus in ASP.NET Core

Design a production agent bus in ASP.NET Core 10 with EF Core, SignalR and PostgreSQL, covering agent registration, topic channels, at least once delivery, claim semantics with row level locking, and self contained hand off payloads.

Multi-agent systems fail in boring ways: two agents claim the same task, a disconnected agent misses its hand-off, a heartbeat proves the TCP socket is open but not that the agent is actually doing anything. Before adding the fifth agent to your workflow, you need a bus with defined delivery guarantees, not an improvised pub/sub over in-memory channels.

This is Part 5 of the series. We now have individual agents that can plan, call tools, and emit results. Here we wire them together with a structured bus: registration, direct messages, topic channels, and task hand offs, all durable in PostgreSQL, with SignalR pushing live notifications to connected agents.

The Model: What the Bus Actually Does

The bus has four responsibilities:

  1. Registration. An agent announces its capabilities (A2A-style Agent Card semantics) and keeps a heartbeat alive.
  2. Direct messages. One agent sends a message to a named agent, guaranteed at-least-once.
  3. Topic channels. Agents subscribe to a topic; all active subscribers receive broadcast messages.
  4. Hand offs. Exactly one agent claims the task for task delegation; the payload is self-contained so a cold agent can resume without additional calls.

SignalR is the notification layer, not the source of truth. Postgres is. If an agent is offline when a message arrives, it polls on reconnect; SignalR just makes the happy path fast.

EF Core Entities

Four entities cover the entire model. Use .NET 10 and EF Core 10.

public class Agent
{
    public Guid AgentId { get; set; }
    public string Name { get; set; } = default!;
    public string[] Capabilities { get; set; } = []; // stored as text[]
    public AgentStatus Status { get; set; } = AgentStatus.Offline;
    public DateTime LastHeartbeatUtc { get; set; }
    public string? SignalRConnectionId { get; set; }
}
 
public enum AgentStatus { Offline, Online, Stale }
 
public class BusMessage
{
    public Guid MessageId { get; set; }
    public Guid? TargetAgentId { get; set; }   // null = channel broadcast
    public string? Channel { get; set; }        // null = direct
    public Guid SenderAgentId { get; set; }
    public string Payload { get; set; } = default!; // JSONB
    public bool Processed { get; set; }
    public DateTime CreatedUtc { get; set; }
}
 
public class ChannelSubscription
{
    public Guid AgentId { get; set; }
    public string Channel { get; set; } = default!;
    public DateTime SubscribedAt { get; set; }
}
 
public class Handoff
{
    public Guid HandoffId { get; set; }
    public Guid OriginAgentId { get; set; }
    public Guid ConversationId { get; set; }
    public string Objective { get; set; } = default!;
    public string MessageHistory { get; set; } = default!; // JSONB array
    public string ToolContext { get; set; } = default!;    // JSONB object
    public int ContextVersion { get; set; }               // optimistic concurrency
    public HandoffStatus Status { get; set; } = HandoffStatus.Unclaimed;
    public Guid? ClaimedByAgentId { get; set; }
    public DateTime? ClaimedAt { get; set; }
    public DateTime CreatedUtc { get; set; }
    public DateTime ExpiresUtc { get; set; }
}
 
public enum HandoffStatus { Unclaimed, Claimed, Completed, Expired }

In OnModelCreating, map MessageHistory and ToolContext as jsonb columns and add an index on Handoff.Status so the claim query stays fast under load.

Claim Semantics: Never Give the Same Task Twice

This is the part that most implementations get wrong. Without a lock, two agents executing SELECT ... WHERE status = 'Unclaimed' LIMIT 1 concurrently will both see the same row and both attempt to claim it.

SELECT FOR UPDATE SKIP LOCKED is the correct tool. EF Core 10 still has no first-party LINQ operator for this, so use raw SQL inside an explicit transaction. The community library EntityFrameworkCore.Locking.PostgreSQL wraps this cleanly, but here is the explicit version so the semantics are transparent:

public async Task<Handoff?> TryClaimHandoffAsync(
    Guid agentId, CancellationToken ct = default)
{
    await using var tx = await _db.Database
        .BeginTransactionAsync(IsolationLevel.ReadCommitted, ct);
 
    // Raw SQL: lock exactly one unclaimed row, skip any already locked by a peer.
    var handoff = await _db.Handoffs
        .FromSqlRaw(
            """
            SELECT * FROM "Handoffs"
            WHERE "Status" = 0             -- Unclaimed
              AND "ExpiresUtc" > NOW()
            ORDER BY "CreatedUtc"
            LIMIT 1
            FOR UPDATE SKIP LOCKED
            """)
        .FirstOrDefaultAsync(ct);
 
    if (handoff is null)
    {
        await tx.RollbackAsync(ct);
        return null;
    }
 
    handoff.Status = HandoffStatus.Claimed;
    handoff.ClaimedByAgentId = agentId;
    handoff.ClaimedAt = DateTime.UtcNow;
    handoff.ContextVersion++;
 
    await _db.SaveChangesAsync(ct);
    await tx.CommitAsync(ct);
 
    return handoff;
}

The lock is held for the length of the transaction, which is milliseconds. Two concurrent callers: the first commits; the second's SKIP LOCKED skips the now-locked row and either claims a different row or returns null. No race condition, no double-claim.

One additional safety net: set ExpiresUtc on creation (e.g., CreatedUtc + 10 minutes). A background IHostedService sweeps expired Claimed rows. If a claiming agent died before finishing, the sweeper resets Status to Unclaimed so another agent can pick it up. This is your at-least-once re-delivery for hand-offs.

At-Least-Once vs At-Most-Once: Draw the Line Clearly

Message type Delivery mode Rationale
Hand-off At-least-once (Postgres + claim) Task loss is not recoverable
Direct message At-least-once (Postgres + poll/push) Business logic depends on receipt
Channel broadcast At-least-once for durable channels Missed steps corrupt workflow
Heartbeat / UI ping At-most-once (SignalR only) Stale liveness signal is harmless

For any at-least-once handler, idempotency is non-negotiable. Track processed MessageId values in the same transaction as the side effect:

public async Task HandleDirectMessageAsync(BusMessage message, CancellationToken ct)
{
    await using var tx = await _db.Database.BeginTransactionAsync(ct);
 
    // Idempotency guard, same transaction as the effect.
    bool alreadyProcessed = await _db.BusMessages
        .AnyAsync(m => m.MessageId == message.MessageId && m.Processed, ct);
 
    if (alreadyProcessed)
    {
        await tx.RollbackAsync(ct);
        return;
    }
 
    // --- Execute domain logic here ---
    await _domainHandler.ExecuteAsync(message.Payload, ct);
 
    message.Processed = true;
    await _db.SaveChangesAsync(ct);
    await tx.CommitAsync(ct);
}

A crash between SaveChanges and the commit causes a retry that hits the guard. A crash after the commit means the effect happened exactly once and the guard permanently prevents re-execution.

The SignalR Hub

The hub is deliberately thin. Its only jobs are delivering push notifications to connected agents and managing connection-to-agent mapping.

[Authorize]
public class AgentBusHub : Hub
{
    private readonly AgentBusDb _db;
 
    public AgentBusHub(AgentBusDb db) => _db = db;
 
    public override async Task OnConnectedAsync()
    {
        var agentId = Guid.Parse(Context.UserIdentifier!);
        var agent = await _db.Agents.FindAsync(agentId);
        if (agent is not null)
        {
            agent.SignalRConnectionId = Context.ConnectionId;
            agent.Status = AgentStatus.Online;
            agent.LastHeartbeatUtc = DateTime.UtcNow;
            await _db.SaveChangesAsync();
        }
 
        // Re-join topic group subscriptions on reconnect.
        var subs = await _db.ChannelSubscriptions
            .Where(s => s.AgentId == agentId)
            .ToListAsync();
        foreach (var sub in subs)
            await Groups.AddToGroupAsync(Context.ConnectionId, sub.Channel);
 
        await base.OnConnectedAsync();
    }
 
    public override async Task OnDisconnectedAsync(Exception? exception)
    {
        var agentId = Guid.Parse(Context.UserIdentifier!);
        var agent = await _db.Agents.FindAsync(agentId);
        if (agent is not null)
        {
            agent.SignalRConnectionId = null;
            agent.Status = AgentStatus.Offline;
            await _db.SaveChangesAsync();
        }
        await base.OnDisconnectedAsync(exception);
    }
 
    public async Task Heartbeat()
    {
        var agentId = Guid.Parse(Context.UserIdentifier!);
        await _db.Agents
            .Where(a => a.AgentId == agentId)
            .ExecuteUpdateAsync(s =>
                s.SetProperty(a => a.LastHeartbeatUtc, DateTime.UtcNow)
                 .SetProperty(a => a.Status, AgentStatus.Online));
    }
 
    public async Task SubscribeToChannel(string channel)
    {
        var agentId = Guid.Parse(Context.UserIdentifier!);
        _db.ChannelSubscriptions.Add(new ChannelSubscription
        {
            AgentId = agentId,
            Channel = channel,
            SubscribedAt = DateTime.UtcNow
        });
        await _db.SaveChangesAsync();
        await Groups.AddToGroupAsync(Context.ConnectionId, channel);
    }
}

Add MessagePack protocol on the server side to reduce wire size for large conversation-history payloads:

builder.Services.AddSignalR()
    .AddMessagePackProtocol();

For multi-pod deployments, add a Redis or EF Core backplane before shipping. A message sent to an agent connected to pod B from a hub call on pod A is silently dropped without it.

Heartbeat Semantics: Liveness, Not Just Connectivity

A heartbeat that fires on a timer proves the timer is running, not that the agent is responsive. The Heartbeat() hub method above is better because the agent has to run application code to call it, but it still does not prove the agent's task loop is healthy.

For a stronger liveness signal, include a WorkerStatus field in the heartbeat payload: current task ID, queue depth, last completed task timestamp. Store it on the Agent entity. A background sweeper (IHostedService) running every 15 seconds marks agents as Stale if LastHeartbeatUtc < UtcNow - 30s, then re-queues any Claimed hand-offs owned by stale agents by resetting their Status to Unclaimed. This is your failure recovery path for crashed agents.

The Hand-off Payload Contract

A cold agent claiming a hand-off must be able to continue the task without any additional API calls. The payload must be self-contained:

{
  "HandoffId": "...",
  "OriginAgentId": "...",
  "ConversationId": "...",
  "Objective": "Summarise Q3 financial results and draft email",
  "MessageHistory": [ ... ],
  "ToolContext": { "availableTools": [...], "lastToolOutput": {...} },
  "ClaimedAt": "2026-09-17T10:42:00Z",
  "ContextVersion": 3
}

ContextVersion enables optimistic concurrency checks if the originating agent tries to update shared state while the claiming agent is executing. MessageHistory carries the full conversation so the claiming agent's LLM call has full context. ToolContext includes the last tool output so the agent doesn't repeat completed work.

What to Watch in Production

Three metrics matter most. First, handoff_claim_latency, the time from CreatedUtc to ClaimedAt. Spikes here mean your agent pool is undersized or agents are blocked. Second, stale_agent_count, where a rising number signals a systemic agent crash loop. Third, message_queue_depth per agent, where unbounded growth means an agent is alive but not processing. Alert on all three before you route real workloads through this bus.

Sources

  1. ASP.NET Core
  2. .NET
  3. @microsoft/signalr - npm
  4. SignalR
  5. core/release-notes/9.0/preview/rc1/aspnetcore.md at main · dotnet/core
  6. NuGet Gallery | Microsoft.AspNetCore.SignalR.Client 10.0.11
  7. ASP.NET Core SignalR New Features — Summary | ABP.IO
  8. ASP.NET Core SignalR clients | Microsoft Learn
Share