
Keeping Your Knowledge Graph Current Without a Dedicated Team
A knowledge graph that nobody updates becomes a lie with good syntax. This article builds the automated maintenance loop — deriving edges from git history, deployment manifests, and OpenTelemetry service maps, reconciling on a schedule, and handling conflicts with provenance on every edge.
The Maintenance Problem Nobody Plans For
You built a knowledge graph to give your AI agent a trustworthy model of your system. Six months later it still describes a service that was decommissioned in March, a dependency that was swapped in a migration nobody documented, and a team that was reorganised. The graph is not wrong in an obvious way. It is coherent, it parses, it answers queries. It is just describing a system that no longer exists.
The fix is not better discipline. It is a maintenance loop that derives edges from what the system already produces, namely git history, deployment manifests and OpenTelemetry trace data, and reconciles those observations against the graph on a schedule. This article builds that loop.
All examples target Neo4j 2026.09.0 (or 5.26 LTS) with Neo4j.Driver 6.2.1 on .NET 8+.
Where the Edges Come From
Three sources give you most of what you need without asking anyone to write documentation:
Git history. CODEOWNERS, directory structure, and cross-repository references encode ownership and coupling. A weekly parse of git log --follow and git blame tells you which team touched which module last, and whether a module has been silent for a year.
Deployment manifests. Kubernetes Deployment and Service manifests, Helm values, and Terraform state files enumerate what is actually running, which images are deployed, and which secrets each workload mounts. These are ground truth. If it is not in the manifest, it is not in the cluster.
OpenTelemetry service maps. The OTel Collector contrib service_graph connector pairs client and server spans, extracting service-to-service call edges as metrics. Configuring dimensions to include http.request.method, http.response.status_code, and rpc.service gives you typed, observable dependency edges that reflect real traffic rather than stale architecture diagrams.
One tuning note: the default store.ttl: 2s means the connector will miss spans where client and server arrive more than two seconds apart. If you have long-running gRPC calls, raise this to at least 10s.
The Reconciliation Worker
The worker runs as a .NET BackgroundService. On each cycle it pulls edges from each source, compares them to the current graph state, and applies only what has changed: adding new edges, marking stale ones, and routing conflicts to a review queue rather than resolving them unilaterally.
public sealed class GraphReconciliationWorker : BackgroundService
{
private readonly IGraphSourceAggregator _sources;
private readonly IGraphRepository _graph;
private readonly IReviewQueue _reviewQueue;
private readonly ILogger<GraphReconciliationWorker> _logger;
private static readonly TimeSpan Interval = TimeSpan.FromHours(4);
public GraphReconciliationWorker(
IGraphSourceAggregator sources,
IGraphRepository graph,
IReviewQueue reviewQueue,
ILogger<GraphReconciliationWorker> logger)
{
_sources = sources;
_graph = graph;
_reviewQueue = reviewQueue;
_logger = logger;
}
protected override async Task ExecuteAsync(CancellationToken ct)
{
while (!ct.IsCancellationRequested)
{
try
{
var observed = await _sources.CollectAsync(ct);
var current = await _graph.LoadEdgeSnapshotAsync(ct);
var diff = EdgeDiff.Compute(current, observed);
await _graph.UpsertEdgesAsync(diff.New, ct);
await _graph.MarkStaleAsync(diff.Missing, ct);
foreach (var conflict in diff.Conflicts)
await _reviewQueue.EnqueueAsync(conflict, ct);
_logger.LogInformation(
"Reconciliation complete. New={New}, Stale={Stale}, Conflicts={Conflicts}",
diff.New.Count, diff.Missing.Count, diff.Conflicts.Count);
}
catch (Exception ex) when (ex is not OperationCanceledException)
{
_logger.LogError(ex, "Reconciliation cycle failed");
}
await Task.Delay(Interval, ct);
}
}
}The critical design choice is MarkStaleAsync rather than delete. Deletion is irreversible and loses provenance. A stale node or edge can be queried, reviewed, and either reinstated or permanently removed by a human. Agents are told to treat stale nodes as unverified, not as absent.
Provenance on Every Edge
Every edge written by the reconciliation worker carries provenance properties so you can answer: who said this, when, and how confident were they?
The AuthentiCity knowledge graph (Zenodo, August 2026), spanning 176.8 GiB of Neo4j store across five cities, models this using confidence-weighted ENRICHED_BY edges and explicit source identifiers on each relationship. That pattern scales.
In Cypher:
private const string UpsertEdgeCypher = """
MERGE (a:Service {id: $sourceId})
MERGE (b:Service {id: $targetId})
MERGE (a)-[r:CALLS]->(b)
ON CREATE SET
r.firstSeen = datetime(),
r.lastConfirmed = datetime(),
r.source = $source,
r.confidence = $confidence,
r.stale = false,
r.provenanceRunId = $runId
ON MATCH SET
r.lastConfirmed = datetime(),
r.source = $source,
r.confidence = $confidence,
r.stale = false,
r.provenanceRunId = $runId
""";
public async Task UpsertEdgesAsync(IReadOnlyList<ObservedEdge> edges, CancellationToken ct)
{
await using var session = _driver.AsyncSession();
foreach (var edge in edges)
{
await session.RunAsync(UpsertEdgeCypher, new
{
sourceId = edge.SourceId,
targetId = edge.TargetId,
source = edge.SourceSystem, // e.g. "otel", "git", "manifest"
confidence = edge.Confidence, // 0.0 to 1.0
runId = edge.RunId
});
}
}Note the version pinning. If you are on any Neo4j release between 2025.11 and 2026.01.3, a verified bug in dynamic relationship types with composite multi-property indexes causes MERGE to silently write to the wrong node or issue CREATE operations that fail silently. Fixed in 2026.01.4 (released 10 February 2026). Run the staleness query below and audit your relationship counts if you are upgrading from that window.
Conflict Handling: When Two Sources Disagree
Sources will contradict each other. OTel says Service A calls Service B. The Kubernetes manifest shows Service B was removed last sprint. Git history shows no recent activity in Service B's repository.
The rule is: the automated pass is not allowed to resolve source disagreements alone. It can log, score, and enqueue. It cannot pick a winner.
The conflict score is a weighted combination of source recency and confidence. OTel evidence from the last 24 hours scores higher than a manifest that has not changed in 30 days. But even a high-confidence automated score is not sufficient. Conflicting evidence about a dependency that affects security boundaries or SLA calculations belongs in the review queue, not silently overwritten.
The review queue is a table in your own store (or a simple Azure Service Bus topic). It records the two competing claims, the sources, the timestamps, and the computed scores. A weekly rotation, one engineer for thirty minutes, clears most of it.
Querying for Stale Nodes
The staleness query is the heartbeat check for graph health. Run it as part of every reconciliation cycle and expose it to the agent so it can self-report uncertainty:
private const string StaleEdgeCypher = """
MATCH (a)-[r]->(b)
WHERE r.stale = true
OR r.lastConfirmed < datetime() - duration('P7D')
RETURN
a.id AS source,
b.id AS target,
type(r) AS relationship,
r.lastConfirmed AS lastConfirmed,
r.source AS evidenceSource
ORDER BY r.lastConfirmed ASC
LIMIT 200
""";
public async Task<IReadOnlyList<StaleEdge>> QueryStaleAsync(CancellationToken ct)
{
await using var session = _driver.AsyncSession();
var result = await session.RunAsync(StaleEdgeCypher);
return await result.ToListAsync(r => new StaleEdge(
Source: r["source"].As<string>(),
Target: r["target"].As<string>(),
Relationship: r["relationship"].As<string>(),
LastConfirmed: r["lastConfirmed"].As<LocalDateTime>(),
EvidenceSource: r["evidenceSource"].As<string>()
), ct);
}Seven days is a reasonable staleness threshold for most systems, long enough to survive a quiet weekend and short enough to catch a decommissioned service before an agent routes a customer query through it.
Honest Limits
The approach works well for what is running and how it is connected. It works poorly for:
Intent and rationale. Git commit messages are inconsistent. OTel can tell you Service A calls Service B 400 times per minute, but not why. That knowledge lives in design documents, ADRs, and conversations. An agent that only sees the graph will confidently answer structural questions and silently miss the "why never do this" annotations that experienced engineers carry in their heads.
Ownership during reorgs. CODEOWNERS files lag by weeks. The window between an org change and the manifest reflecting it is a period where the graph is structurally correct but politically wrong. Agents that route escalations using ownership edges during that window will get it wrong.
Security-sensitive edges. Which service has read access to which secret store is derivable from manifests, but the intended access policy versus the actual policy can diverge. Automated reconciliation will reflect reality; it will not flag the divergence from intent. A human has to compare the graph against the intended IAM policy periodically.
The reconciliation worker is not a replacement for architectural review. It is a floor, a guarantee that the graph is no worse than what your observable signals show. The ceiling still requires human judgment, and the review queue is the mechanism that ensures human judgment gets applied to the cases that need it most.
Practical Checklist Before You Ship
- Confirm Neo4j version is 2026.01.4 or later if you use composite indexes on relationship types.
- Use Neo4j.Driver 6.2.1. The 4.4 driver is security-only with no new feature backports.
- Tune
service_graphconnectorstore.ttlabove2sif you have long-running RPC calls. - Verify which APOC procedures you depend on are in Core vs Extended, since they are separate downloads since Neo4j 5.0.
- Every automated edge write must carry
source,confidence,lastConfirmed, andprovenanceRunId. - The review queue is not optional. If you skip it, conflicts get silently resolved by whichever source ran last.
Sources
- Neo4j deployment center
- Neo4j
- Neo4j
- GitHub - neo4j-contrib/neo4j-apoc-procedures: Awesome Procedures On Cypher for Neo4j - codenamed "apoc" If you like it, please ★ above ⇧
- GitHub - neo4j/apoc · GitHub
- AuthentiCity: A Multi-Source Provenance-Aware Knowledge Graph and Benchmark for 3D City Models
- Neo4j APOC Procedures: The Definitive Guide
- Releases · neo4j-contrib/neo4j-apoc-procedures
Keep reading

October 6, 2026 · 7 min
Hybrid Retrieval in C#: Combining Vector Search, Graph Traversal, and Reranking for Agent Memory
Vector search finds candidate entry points; graph traversal expands context around them; a cross-encoder reranker decides what actually reaches the model. This article implements that full pipeline in C# with real token budgets, RRF fusion, and deduplication.
Read
October 5, 2026 · 7 min
Building an Agentic System in .NET, Part 6: Redaction, Audit and the Safety Layer
If you archive every agent session, you also archive every secret anyone ever pasted into one. This last part covers the full safety layer: redaction pipelines, audit trails, prompt injection threat modelling and retention pruning, with code.
Read
September 28, 2026 · 7 min
Graph-Native Data Structures in C#, Part 6: DAGs and Dependency Graphs
Model a build system dependency graph in NebulaGraph, keep it acyclic from C# at insert time, get the execution order with Kahn's algorithm, and answer what is blocked right now. With real nGQL and C# code.
Read