All articles
AI/October 5, 2026/7 min read

Building an Agentic System in .NET, Part 6: Redaction, Audit and the Safety Layer

If you archive every agent session, you also archive every secret anyone ever pasted into one. This last part covers the full safety layer: redaction pipelines, audit trails, prompt injection threat modelling and retention pruning, with code.

Prompt injection is OWASP's LLM01:2025, the top vulnerability in LLM based systems, and yet most agentic .NET projects I review have no redaction pipeline at all. They have a SessionStore, they have JsonSerializer, and they have a lot of optimism. This article closes the series by building the safety layer that the first five parts deliberately deferred: redaction before storage, tamper-evident audit, per-workspace ownership, and the threat model of a tool that can write.

Why Redaction Must Run Before Storage

The 70% failure rate of visual-blackout redaction exists because teams treat redaction as a presentation concern. It is not. It is a data concern. Once a token, API key, or connection string reaches your database, you need a migration, a key rotation, and a breach notification conversation. The only safe moment is the pipeline stage before the INSERT.

The pattern: every agent turn produces a ConversationTurn object. Before that object is handed to the store, it passes through a RedactionPipeline that runs detectors in order, deterministic regex first, entropy heuristics second, contextual model third, and replaces sensitive spans with reversible tokens stored separately in a masked-value vault.

// RedactionPipeline.cs
public sealed class RedactionPipeline
{
    private readonly IEnumerable<IRedactionDetector> _detectors;
    private readonly IMaskedValueVault _vault;
    private readonly IFlaggedItemQueue _reviewQueue;
 
    public RedactionPipeline(
        IEnumerable<IRedactionDetector> detectors,
        IMaskedValueVault vault,
        IFlaggedItemQueue reviewQueue)
    {
        _detectors = detectors;
        _vault = vault;
        _reviewQueue = reviewQueue;
    }
 
    public async Task<RedactionResult> ProcessAsync(
        string workspaceId, string turnId, string text,
        CancellationToken ct = default)
    {
        var spans = new List<DetectedSpan>();
        foreach (var detector in _detectors)
            spans.AddRange(await detector.DetectAsync(text, ct));
 
        // Merge overlapping spans, highest-confidence wins
        var merged = MergeSpans(spans);
        var masked = text;
        var mappings = new List<MaskMapping>();
 
        foreach (var span in merged.OrderByDescending(s => s.Start))
        {
            var raw = text[span.Start..span.End];
            var token = $"[REDACTED:{span.Category}:{Guid.NewGuid():N}]";
            var vaultKey = await _vault.StoreAsync(workspaceId, raw, ct);
            mappings.Add(new MaskMapping(token, vaultKey, span));
            masked = masked[..span.Start] + token + masked[span.End..];
 
            if (span.Confidence < 0.80)
                await _reviewQueue.EnqueueAsync(
                    new FlaggedItem(workspaceId, turnId, token, span), ct);
        }
 
        return new RedactionResult(masked, mappings);
    }
 
    private static IReadOnlyList<DetectedSpan> MergeSpans(
        IEnumerable<DetectedSpan> spans)
    {
        var sorted = spans.OrderBy(s => s.Start).ToList();
        var result = new List<DetectedSpan>();
        foreach (var span in sorted)
        {
            if (result.Count > 0 && span.Start < result[^1].End)
            {
                var last = result[^1];
                result[^1] = last with
                {
                    End = Math.Max(last.End, span.End),
                    Confidence = Math.Max(last.Confidence, span.Confidence)
                };
            }
            else result.Add(span);
        }
        return result;
    }
}

Detector Implementations

The three tiers that matter in practice:

Tier 1, regex detectors. for connection strings (Server=.*;Password=), JWT patterns, AWS key prefixes (AKIA[0-9A-Z]{16}), and GitHub PATs (ghp_[A-Za-z0-9]{36}). These are zero-latency and zero-false-negative for structured secrets.

Tier 2, Shannon entropy detector.: any contiguous token above 4.5 bits/character and longer than 20 characters that is not a URL or a GUID is a candidate. This catches random-looking API keys that don't match a known prefix.

Tier 3, Microsoft Presidio, through a REST sidecar or the Python SDK.: covers names, addresses, NI numbers, IBAN, passport numbers. Keep Presidio as a sidecar. Its latency is fine for async storage paths but not for real-time streaming. For unstructured free text, the Azure Conversation PII API (2026-11-15-preview, model 2026-04-15-preview) handles multi-turn conversation structures and is worth evaluating if you're already in the Azure footprint.

One important disclaimer: GDPR Article 4(5) defines pseudonymization as leaving a data subject re-identifiable when the additional information (your vault) is available. Replacing Alice Smith with PERSON_001 does not move the data outside GDPR. Rare job title + small town + event date can re-identify without any name at all. Treat the complete turn context as the assessment unit.

Unit Testing the Pipeline

Test the pipeline with property-based inputs, not just happy-path strings:

[Theory]
[InlineData("Server=myserver;Password=hunter2;", "ConnectionString")]
[InlineData("AKIAIOSFODNN7EXAMPLE secret", "AwsKey")]
[InlineData("Bearer eyJhbGciOiJSUzI1NiJ9.abc.def", "Jwt")]
public async Task Pipeline_RedactsKnownPatterns(string input, string expectedCategory)
{
    var vault = new InMemoryMaskedValueVault();
    var queue = new InMemoryFlaggedItemQueue();
    var pipeline = new RedactionPipeline(
        new IRedactionDetector[] { new RegexRedactionDetector() },
        vault, queue);
 
    var result = await pipeline.ProcessAsync("ws-001", "turn-001", input);
 
    Assert.DoesNotContain("hunter2", result.MaskedText);
    Assert.DoesNotContain("AKIAIOSFODNN7EXAMPLE", result.MaskedText);
    Assert.Contains(expectedCategory, result.MaskedText);
    Assert.NotEmpty(vault.All());
}
 
[Fact]
public async Task Pipeline_SendsLowConfidenceToReviewQueue()
{
    var vault = new InMemoryMaskedValueVault();
    var queue = new InMemoryFlaggedItemQueue();
    var entropy = new EntropyRedactionDetector(threshold: 4.5, minLength: 20,
        baseConfidence: 0.65); // below 0.80 threshold
    var pipeline = new RedactionPipeline(
        new IRedactionDetector[] { entropy }, vault, queue);
 
    await pipeline.ProcessAsync("ws-001", "turn-002",
        "config value: xK9mP2qRvL8nW3yT6jA1cB5eH0uM4sZ7");
 
    Assert.Single(queue.All());
}

Flagged-Item Review Endpoint

Items below the confidence threshold land in a review queue. Operators confirm or dismiss them via a minimal admin endpoint. Confirmed items get re-stored with elevated confidence and marked Verified; dismissed items are unmasked in the vault and the token is replaced in storage:

// FlaggedItemsEndpoint.cs, registered with app.MapGroup("/admin/flagged")
app.MapGet("/", async (IFlaggedItemQueue queue, string workspaceId) =>
    Results.Ok(await queue.GetPendingAsync(workspaceId)));
 
app.MapPost("/{id}/confirm", async (
    string id, IFlaggedItemQueue queue, IMaskedValueVault vault) =>
{
    var item = await queue.GetAsync(id);
    if (item is null) return Results.NotFound();
    await vault.MarkVerifiedAsync(item.VaultKey);
    await queue.ResolveAsync(id, resolution: FlagResolution.Confirmed);
    return Results.Ok();
});
 
app.MapPost("/{id}/dismiss", async (
    string id, IFlaggedItemQueue queue, IMaskedValueVault vault,
    ISessionStore store) =>
{
    var item = await queue.GetAsync(id);
    if (item is null) return Results.NotFound();
    var original = await vault.RetrieveAsync(item.WorkspaceId, item.VaultKey);
    await store.RestoreTokenAsync(item.TurnId, item.MaskToken, original);
    await vault.DeleteAsync(item.WorkspaceId, item.VaultKey);
    await queue.ResolveAsync(id, resolution: FlagResolution.Dismissed);
    return Results.Ok();
});

Audit Trails and Prompt-Injection Threat Model

ATR-2026-00550 (published 28 May 2026, severity: critical) defines the canonical indirect prompt-injection trace shape: an untrusted RETRIEVER span followed by a privileged TOOL span. EchoLeak (CVE-2025-32711) demonstrated this path in production against Microsoft 365 Copilot, with no credentials, no malware and no user interaction. Anthropic's official Git MCP server shipped three exploitable injection CVEs.

The audit trail's job is to make this trace shape queryable after the fact. Every tool call must emit a structured audit event with:

  • WorkspaceId and AgentSessionId
  • ToolName, ToolInputHash (not the raw input, because the input may contain the payload)
  • SourceSpan: was the input retrieved from an untrusted source?
  • PrivilegeLevel: Read, Write or Admin
  • OutcomeStatus and OutcomeHash

Store audit events in append-only storage. In PostgreSQL this means an insert-only table with a trigger that prevents UPDATE and DELETE, plus a nightly hash-chain confirmation job. In Azure, an immutable blob container with a time-based policy achieves the same tamper-evidence property. Without deterministic replay, you reconstruct incidents from memory; with it, you run a query.

For privileged write actions specifically, apply the invariant from the research: every published control has been bypassed in isolation. Plan for five to seven independent layers: untrusted content fencing, instruction-hierarchy prompts, output filters, human-in-the-loop confirmation, tool-level allow-listing, audit alerting, and a kill-switch circuit breaker. None of these in isolation is sufficient.

Per-Workspace Ownership and Retention Pruning

Multi tenant isolation is not a future concern, it is a concern now, because the moment you add a second user the data already exists. A real-world breach (8 months undetected) was caused by a session-cookie authentication bug that allowed one tenant to read another's conversations. Row-Level Security at the database layer is non-negotiable: every query against AgentSessions, ConversationTurns, and MaskedValueVault must include a workspace_id predicate enforced by RLS policy, not just application code.

Retention tiers aligned with 2026 guidance:

Context Retention
No identifiable PII 90 days
Email only, not linked to account 30 days
Support ticket conversation 365 days or ticket resolution +90 d
Health / financial data Per HIPAA / PCI-DSS

Automate pruning. The 2026 audit finding against Nakama was exactly that transcripts persisted forever unless manually purged, which GDPR Art. 5(1)(e) does not allow. A background RetentionPruningJob runs nightly:

public sealed class RetentionPruningJob(ISessionStore store, ILogger<RetentionPruningJob> log)
    : BackgroundService
{
    protected override async Task ExecuteAsync(CancellationToken ct)
    {
        while (!ct.IsCancellationRequested)
        {
            await Task.Delay(TimeSpan.FromHours(24), ct);
            var cutoffs = new Dictionary<RetentionTier, DateTimeOffset>
            {
                [RetentionTier.NoPii]        = DateTimeOffset.UtcNow.AddDays(-90),
                [RetentionTier.EmailOnly]    = DateTimeOffset.UtcNow.AddDays(-30),
                [RetentionTier.SupportTicket]= DateTimeOffset.UtcNow.AddDays(-365),
            };
 
            foreach (var (tier, cutoff) in cutoffs)
            {
                var pruned = await store.SoftDeleteExpiredAsync(tier, cutoff, ct);
                log.LogInformation(
                    "RetentionPruning: tier={Tier} pruned={Count} cutoff={Cutoff:O}",
                    tier, pruned, cutoff);
            }
 
            // Hard-delete soft-deleted records beyond the 30-day quarantine grace period
            var hardDeleted = await store.HardDeleteQuarantinedAsync(
                DateTimeOffset.UtcNow.AddDays(-30), ct);
            log.LogInformation(
                "RetentionPruning: hard-deleted {Count} quarantined records", hardDeleted);
        }
    }
}

The 30-day soft-delete quarantine handles legal holds: a hold flag on the record suspends hard deletion without requiring you to change the pruning job.

Full Architecture and What I Would Build Differently

The complete safety layer sits as a cross-cutting concern across the agent pipeline: Inbound text → RedactionPipeline → Session Store (masked) + Vault (raw, encrypted at rest, RLS-gated) → Agent Execution → Tool calls emit AuditEvents (append-only) → outbound text through RedactionPipeline again → Response. The RetentionPruningJob runs out-of-band against the Session Store. The review endpoint gives operators a window into the Flagged Item Queue.

If I were starting this series again, three things would change:

First, I would make the WorkspaceId a first-class type from Part 1, not a string. Leaking it across workspace boundaries at the database layer is the most common multi-tenant mistake and the strongly-typed wrapper would have made the compiler catch it earlier.

Second, I would add trace shape alerting, the untrusted retriever into privileged write pattern, from the first tool integration, not as a retrofit. The audit schema needs the SourceSpan.IsTrusted flag from day one; adding it later means backfilling every historic event or accepting a gap.

Third, I would treat the redaction pipeline as a testable, deployable component in its own right with its own integration tests against real Presidio and the Azure Conversation PII API, not as something wired into the agent and tested only through end-to-end tests. The $4.88M average breach cost and the 2025 enforcement actions make this worth the investment in a standalone test harness.

The average cost of a PII breach in 2025 reached $4.88 million. A 2025 ed-tech enforcement action cost $5.1 million specifically for inadequate deletion controls. The Polish bank enforcement found no breach and no bad actor, just data collected beyond what the processing purpose required. The compliance floor for agent systems is now the same as for any other personal-data processor. Build the safety layer first.

Sources

  1. Conversation Personally Identifiable Information (PII) redaction overview - Foundry Tools | Microsoft Learn
  2. Identify and extract Personally Identifiable Information (PII) from conversations - Foundry Tools | Microsoft Learn
  3. The Complete Guide to PII Redaction in 2026 | Redactable
  4. Pimloc
  5. Microsoft's Semantic Kernel - The Cracked Kernel - Nuka-AI Research Series
  6. When prompts become shells: RCE vulnerabilities in AI agent frameworks | Microsoft Security Blog
  7. Gateway-Level PII Redaction Before Provider Transmission
  8. What are the security advantages of using Semantic Kernel over other open-source AI frameworks? | iCertGlobal Community
Share