All articles
Databases/July 16, 2026/6 min read

Graph-Native Data Structures in C#, Part 4: Graphs and Adjacency with a Social Follow Network

Part 4 of the series models a directed social follow network in NebulaGraph, then surfaces neighbour queries, mutual connections, friend-of-friend suggestions, and super-node protection as typed C# async methods.

If you have modelled a tree or a list, that was the warm up. The reason graph databases exist is adjacency at scale, the kind of traversal that makes a relational database suffer once you are three joins deep. Social networks are the classic case, and follow relationships are graph thinking at its most basic: directed edges, variable depth hops, set intersections, and the super node sitting behind every celebrity account.

This is Part 4 of the series. Part 1 set up the BaseGraphRepository, the NebulaNet connection pool, and the VID convention 'type:id'. The same foundations apply here. We swap the movie and actor schema for user and follows, but the structure is the same.

All the nGQL here targets NebulaGraph 3.x Community Edition. On Enterprise v5.x you can use the ISO/IEC 39075 GQL MATCH equivalents, and the logic carries over directly.

Modelling the Follow Network

Schema

CREATE SPACE IF NOT EXISTS socialnet(vid_type=FIXED_STRING(30));
USE socialnet;
 
CREATE TAG IF NOT EXISTS user(handle STRING, name STRING);
CREATE EDGE IF NOT EXISTS follows(since INT);

Vertices carry a user tag. The VID convention is 'user:' + handle, so 'user:ada', 'user:bob' and so on. Keep handles short enough that the whole VID fits in FIXED_STRING(30). The prefix takes five characters, which leaves 25 for the handle. Edge direction is physical: follows goes from follower to followee.

Directed or Symmetric: Follows Versus Friends

follows is asymmetric by nature. Alice follows Bob, Bob may not follow back. One directed edge covers that.

Friends, meaning a mutual follow, is harder. There are two ways to do it.

Option A, two edges by convention. When a friendship forms, insert (alice)-[:follows]->(bob) and (bob)-[:follows]->(alice). Reads stay simple, because GO FROM 'user:alice' OVER follows gives you everyone Alice follows and Bob is in there through his own edge. The cost is twice the edge storage and two writes per friendship.

Option B, one canonical edge plus BIDIRECT. Store one direction, decided by a rule such as lower ID to higher ID, and always query with BIDIRECT. Writes are cheap. Reads need discipline, because everyone on the team has to remember BIDIRECT, and a plain GO OVER follows turns into a correctness bug that says nothing.

In a follow network where "following" and "friends" really are different things, Option A is almost always right. The extra storage is cheap. Clear semantics are not.

Typed C# Queries: Followers and Following

The repository wrapper matches earlier parts of the series. One gotcha with nebula-net: session.Release() is not IAsyncDisposable, so wrap every session in try/finally.

public sealed class FollowRepository : BaseGraphRepository
{
    public FollowRepository(NebulaPool pool) : base(pool) { }
 
    // Who does 'handle' follow?
    public async Task<List<string>> GetFollowingAsync(string handle)
    {
        var vid = $"'user:{handle}'";
        var nGql = $"""
            GO FROM {vid} OVER follows
            YIELD dst(edge) AS followeeVid
            """;
        return await ExecuteScalarListAsync<string>(nGql, "followeeVid");
    }
 
    // Who follows 'handle'?
    public async Task<List<string>> GetFollowersAsync(string handle)
    {
        var vid = $"'user:{handle}'";
        var nGql = $"""
            GO FROM {vid} OVER follows REVERSELY
            YIELD dst(edge) AS followerVid
            """;
        return await ExecuteScalarListAsync<string>(nGql, "followerVid");
    }
 
    // Out-degree: number of accounts 'handle' follows
    public async Task<long> GetFollowingCountAsync(string handle)
    {
        var vid = $"'user:{handle}'";
        var nGql = $"""
            GO FROM {vid} OVER follows
            YIELD count(*) AS degree
            """;
        return await ExecuteScalarAsync<long>(nGql, "degree");
    }
}

REVERSELY flips the traversal to incoming edges, which is how you get followers without storing a second edge type. count(*) counts edges traversed, which is the out degree as long as each pair has at most one follows edge.

BFS, DFS, and Why You Always Bound the Depth

GO does a walk, not a simple path traversal, so it can re enter a vertex on a cyclic graph. Run GO FROM 'user:ada' OVER follows with no step bound on a graph that has mutual follows and you have an endless traversal waiting for you.

Bound the depth every time:

GO 3 STEPS FROM 'user:ada' OVER follows YIELD DISTINCT dst(edge)

For open ended exploration you can use MATCH with a range, MATCH p=(v)-[:follows*1..3]->(w). The upper bound is not optional. Leave it out and the engine either refuses to run the query or, on older versions, runs until it runs out of memory.

NebulaGraph Enterprise v5.2 claims 100 times faster path queries thanks to its in database compute engine, but the rule is the same on every version. An unbounded traversal on a real social graph with millions of edges is a denial of service attack on your own database.

For recommendations and social features, anything past depth 3 rarely gives you a useful signal and often gives you a result set too big to do anything with. Use 2 for suggestions. Allow 3 only with a strict LIMIT and a SAMPLE clause.

Practical Queries: Mutual Connections and Friend Suggestions

public async Task<List<string>> GetMutualAsync(string handleA, string handleB)
{
    // Intersection of followees: people both A and B follow
    var vidA = $"'user:{handleA}'";
    var vidB = $"'user:{handleB}'";
    var nGql = $"""
        $a = GO FROM {vidA} OVER follows YIELD dst(edge) AS vid;
        $b = GO FROM {vidB} OVER follows YIELD dst(edge) AS vid;
        YIELD $a.vid AS mutual
        WHERE $a.vid IN $b.vid
        """;
    return await ExecuteScalarListAsync<string>(nGql, "mutual");
}
 
public async Task<List<string>> SuggestFollowsAsync(string handle, int limit = 20)
{
    // Classic 2-hop: who do my followees follow, that I don't already follow?
    var vid = $"'user:{handle}'";
    var nGql = $"""
        $already = GO FROM {vid} OVER follows YIELD dst(edge) AS vid;
        GO 2 STEPS FROM {vid} OVER follows
        YIELD DISTINCT dst(edge) AS candidate
        WHERE candidate != {vid}
          AND candidate NOT IN $already.vid
        | LIMIT {limit}
        """;
    return await ExecuteScalarListAsync<string>(nGql, "candidate");
}

Three things are worth pointing out:

  • YIELD DISTINCT on the two hop query is not optional. Several intermediate users may follow the same third person, and without DISTINCT that person appears once per path, which wrecks any ranking you do afterwards.
  • The $already pipe holds the current followees so you can exclude them. Without that filter the most popular accounts near the user, which they almost certainly follow already, take over every suggestion list.
  • WHERE candidate != {vid} stops the user being suggested to themselves through a mutual follow cycle.

Handling Super Nodes with SAMPLE

Every social graph has celebrities, accounts with hundreds of thousands or millions of followers. Walking every incoming or outgoing edge on one of those per hop is a guaranteed latency spike.

SAMPLE caps how many edges are explored per hop:

GO 2 STEPS FROM 'user:ada' OVER follows
YIELD DISTINCT dst(edge) AS candidate
SAMPLE [100, 50]

[100, 50] means at most 100 edges at hop one and at most 50 at hop two. The list length must match the hop count, so GO 2 STEPS needs exactly two values. Getting that wrong is a runtime error, not a compile time one, so validate it in your repository wrapper.

public async Task<List<string>> SuggestFollowsSafeAsync(
    string handle, int limit = 20, int[] samplePerHop = null)
{
    samplePerHop ??= [100, 50];
    if (samplePerHop.Length != 2)
        throw new ArgumentException(
            "samplePerHop must have exactly 2 values for a 2-hop traversal.",
            nameof(samplePerHop));
 
    var vid = $"'user:{handle}'";
    var sampleClause = $"SAMPLE [{samplePerHop[0]}, {samplePerHop[1]}]";
    var nGql = $"""
        $already = GO FROM {vid} OVER follows YIELD dst(edge) AS vid;
        GO 2 STEPS FROM {vid} OVER follows
        YIELD DISTINCT dst(edge) AS candidate
        WHERE candidate != {vid}
          AND candidate NOT IN $already.vid
        {sampleClause}
        | LIMIT {limit}
        """;
    return await ExecuteScalarListAsync<string>(nGql, "candidate");
}

Sampling adds statistical bias, because you do not see every two hop candidate. For a suggestion feature that is fine and often better. Showing a rotating sample of second degree neighbours feels more alive than an exhaustively ranked list, and it keeps p99 latency steady no matter who is in the path.

Reachability Check

A lighter version of the same idea: can A reach B within N hops at all? This is what powers "you and Bob have a second degree connection".

GO 1 TO 3 STEPS FROM 'user:ada' OVER follows
YIELD dst(edge) AS reached
| WHERE $-.reached == 'user:bob'
| LIMIT 1

GO 1 TO 3 STEPS explores hops one through three in one query, and with LIMIT 1 it stops at the first match. In C#, turn that into a Task<bool> by checking whether the result list is empty.

What Is Next

Part 5 moves from traversal to ranking: putting weight properties on edges and using them to surface quality adjusted recommendations instead of raw hop count proximity. The follows edge already carries since. We add an interaction_score and show how filtering on edge properties changes the query strategy quite a lot.

The SuggestFollowsAsync method built here is where that ranking layer starts.

Sources

  1. GO - Nebula Graph Database Manual
  2. GO - NebulaGraph Database Manual
  3. nGQL cheatsheet - NebulaGraph Database Manual
  4. nGQL Overview - Nebula Graph Database Manual
  5. Step 5 Use nGQL (CRUD) - NebulaGraph Database Manual
  6. SQL & nGQL - Nebula Graph Database Manual
  7. Gremlin & nGQL - Nebula Graph Database Manual
  8. NebulaGraph Query Language (nGQL)
Share