All articles
Architecture/June 26, 2026/6 min read

Two Bugs Hiding in Our OpenTelemetry Pipeline: CORS Preflights and the NUL Byte That Killed a Whole Batch

Two silent production bugs in an OTLP observability pipeline. One blocked every browser telemetry client, the other dropped whole log batches. Here is the root cause, the fix and the repro for each.

The Setup and the Silence

On paper our observability stack is simple. Clients ship OTLP over HTTP to a gateway on port 4318, the gateway validates, routes and forwards to an ingestion service, and the ingestion service bulk inserts into PostgreSQL. The .NET side is the OpenTelemetry SDK 1.16.0 with OpenTelemetry.Exporter.OpenTelemetryProtocol 1.16.0, ASP.NET Core 9 for both the gateway and the ingestion service, and Npgsql 9.x.

For months everything looked fine. Then two new kinds of client appeared. One was a browser SPA shipping traces with @opentelemetry/sdk-trace-web. The other was a native application that sometimes put a NUL byte inside a log message. Both failed quietly. No alert fired, nothing showed up in the dashboard, the data was just missing.

That is the worst kind of observability bug. The pipeline whose job is to tell you something is broken is itself broken, and it does not tell you.

Background: Why There Is a Gateway

In a multi tenant OTLP deployment you do not put the ingestion service on the internet. The gateway handles CORS, which browser clients need, routes /v1/traces, /v1/metrics, /v1/logs and /v1/profiles to the right backend, and keeps anonymous ingest separate from credentialed endpoints such as the dashboard API or the SignalR hub used for live tail. The ingestion service only ever sees internal traffic, so it does not have to think about CORS at all.

That separation is good architecture. It also creates two boundary layers where a wrong assumption can swallow traffic without a sound.

Bug 1: The CORS Preflight That Blocked Browser Clients

Symptom

Browser clients on a different origin (https://app.example.com) sent zero traces. The network tab showed 204 No Content for the OPTIONS preflight, which looks like success, but there was no Access-Control-Allow-Origin header on it. The browser read that as a denied preflight and never sent the real POST /v1/logs.

Root Cause

The gateway had one blanket CORS policy:

builder.Services.AddCors(options =>
{
    options.AddDefaultPolicy(policy =>
    {
        policy
            .AllowAnyOrigin()
            .AllowAnyMethod()
            .AllowAnyHeader()
            .AllowCredentials(); // <-- the problem
    });
});

The CORS specification does not allow Access-Control-Allow-Origin: * together with Access-Control-Allow-Credentials: true. ASP.NET Core enforces that. When it sees AllowAnyOrigin() and AllowCredentials() together, it drops the Access-Control-Allow-Origin header rather than send an invalid response. So the preflight comes back 204 with no CORS headers at all, and to anything that is not a browser that looks exactly like success.

The policy was originally written for the SignalR hub, which really does need credentials. It got copied onto the OTLP routes, and nobody noticed the wildcard plus credentials trap.

OTLP /v1/logs posts Content-Type: application/x-protobuf, which is not a simple content type, so a browser client always triggers a preflight. There is no way around that.

The Fix

A small piece of terminal middleware that catches OPTIONS on /v1/* before the CORS middleware runs and answers with the right anonymous CORS headers. Browser sourced OTLP does not need credentials. It needs Allow-Origin: *.

// In Program.cs, registered BEFORE app.UseCors() and app.UseAuthentication()
app.Use(async (context, next) =>
{
    if (context.Request.Method == HttpMethods.Options
        && context.Request.Path.StartsWithSegments("/v1"))
    {
        var response = context.Response;
 
        // Only stamp the header if the ingestion service hasn't already set one
        // (protects diagnosability of 5xx responses that pass through)
        if (!response.Headers.ContainsKey("Access-Control-Allow-Origin"))
        {
            response.Headers["Access-Control-Allow-Origin"] = "*";
        }
 
        var requestedHeaders = context.Request.Headers["Access-Control-Request-Headers"].ToString();
        if (!string.IsNullOrEmpty(requestedHeaders))
        {
            response.Headers["Access-Control-Allow-Headers"] = requestedHeaders;
        }
 
        response.Headers["Access-Control-Allow-Methods"] = "POST, OPTIONS";
        response.Headers["Access-Control-Max-Age"] = "86400"; // 24 h, the browser caches the preflight
        response.StatusCode = StatusCodes.Status204NoContent;
        return;
    }
 
    await next(context);
});
 
// The credentialed policy stays on /api and /hubs only
app.UseCors("CredentialedPolicy"); // SignalR, dashboard API

The /api and /hubs routes keep their named credentialed policy with explicit origins. The OTLP routes never touch it. Two different auth boundaries, two different policies. That is the real lesson here.

Bug 2: The NUL Byte That Killed a Whole Batch

Symptom

Intermittent HTTP 500 from the ingestion service, always in bursts. The Postgres log said ERROR: invalid byte sequence for encoding "UTF8": 0x00, SQLSTATE 22021. A native application was serialising log messages that sometimes carried a NUL byte (0x00) in the body, most likely from a C style string boundary in its serialisation layer.

The ingestion service does a bulk INSERT for the whole batch, so one bad record rolled back the entire statement. That is normal SQL transaction behaviour, there is no partial success. From the client's point of view it got a 500, and depending on the retry configuration it either dropped the batch or retried the whole thing, bad record included, in a loop.

Root Cause

PostgreSQL's text type is UTF-8 and cannot store a NUL byte, even though NUL is a valid Unicode code point and legal in other encodings. Every other control character, U+0001 to U+001F, goes in without complaint. Only 0x00 is refused. And because the insert is one statement, that single record takes the whole batch with it.

More fields can carry this than you would expect. In an OTLP log record it is the body, attribute keys, attribute values, nested kvlist keys and values, and severityText. All of them hold arbitrary strings that came from a user.

The Fix

A small sanitiser that removes only NUL bytes, applied at the ingestion boundary before the bulk insert is built:

public static class OtlpSanitizer
{
    /// <summary>
    /// Strips NUL (0x00) from a string. PostgreSQL text columns reject NUL;
    /// all other Unicode control characters are stored without error.
    /// Returns the original reference if no NUL is present (zero allocation).
    /// </summary>
    public static string StripNul(string? value)
    {
        if (value is null) return string.Empty;
        if (!value.Contains('\0')) return value; // fast path, no allocation
        return value.Replace("\0", string.Empty);
    }
 
    /// <summary>
    /// Applies StripNul to all string fields in a decoded OTLP log record
    /// before it is handed to the bulk-insert command builder.
    /// </summary>
    public static void SanitizeLogRecord(LogRecord record)
    {
        record.Body = StripNul(record.Body);
        record.SeverityText = StripNul(record.SeverityText);
 
        foreach (var attr in record.Attributes)
        {
            attr.Key = StripNul(attr.Key);
            if (attr.Value?.StringValue is not null)
                attr.Value.StringValue = StripNul(attr.Value.StringValue);
 
            // Recurse into kvlist values
            if (attr.Value?.KvlistValue is not null)
            {
                foreach (var kv in attr.Value.KvlistValue.Values)
                {
                    kv.Key = StripNul(kv.Key);
                    if (kv.Value?.StringValue is not null)
                        kv.Value.StringValue = StripNul(kv.Value.StringValue);
                }
            }
        }
    }
}

The Contains('\0') check means almost every record, the ones with no NUL, allocates nothing. Only the poisoned ones pay for the Replace.

SanitizeLogRecord runs once per record right after protobuf deserialisation, before the batch command is built. The batch then inserts cleanly. If a caller's data had a NUL they get a 200 and their record is stored with the NUL removed, instead of losing the whole batch.

Verifying the Fixes

Both fixes are small: two files changed, no new dependencies, clean build.

Repro for bug 1, the CORS preflight:

curl -i -X OPTIONS https://gateway.example.com/v1/logs \
  -H "Origin: https://app.example.com" \
  -H "Access-Control-Request-Method: POST" \
  -H "Access-Control-Request-Headers: Content-Type"

Before the fix: 204 No Content with no Access-Control-Allow-Origin header. After: 204 No Content with Access-Control-Allow-Origin: *, Access-Control-Allow-Methods: POST, OPTIONS and Access-Control-Max-Age: 86400.

Repro for bug 2, the NUL byte batch:

Build a minimal OTLP log export request in protobuf, or use a test client, with one record whose body is "hello\x00world". Before the fix: HTTP 500 and the whole batch is gone. After: HTTP 200 and the record is stored as "helloworld", NUL removed, rest of the batch intact.

Both can be run as integration tests against a local Postgres with no mocking.

Takeaways

The two bugs have the same shape. Both were invisible in normal operation. Both came from a client that sat just outside the original design assumptions, a browser SPA and a native app with C style strings. And both got through because a boundary layer, the gateway in one case and the ingestion service in the other, was not defensive.

The CORS bug is a policy layering trap. Credentialed CORS and wildcard origins cannot coexist, by specification. If your gateway serves browser facing OTLP and credentialed API endpoints, they need separate named policies on separate route groups. Copying one default policy over both will quietly kill one of them.

The NUL bug is a data hygiene trap. PostgreSQL's text type has one constraint beyond UTF-8 validity, and it is not one you meet in ordinary application traffic. When you ingest telemetry from clients you do not control, native apps, embedded systems, anything that serialises raw memory, you will meet it eventually. The fix costs almost nothing. Not having it costs you a silent full batch drop every time it fires.

If you run an OTLP gateway in .NET, check two things today. Send an OPTIONS /v1/logs from a cross origin curl and confirm you get Access-Control-Allow-Origin: * back. Then send a log record with a NUL in the body and confirm you get a 200 and not a 500. If either one fails, you have the same bugs we had.

Sources

  1. Releases · open-telemetry/opentelemetry-collector
  2. Releases · open-telemetry/opentelemetry-collector-contrib
  3. Exporters | OpenTelemetry
  4. Releases · open-telemetry/opentelemetry.io
  5. OTLP Exporter Configuration | OpenTelemetry
  6. opentelemetry-otlp 0.32.0 - Docs.rs
  7. opentelemetry-exporter-otlp · PyPI
  8. OTLP Exporter for OpenTelemetry .NET
Share