Back to blog

Gated Content on Optimizely Graph, Part 2: Keeping Graph in Step with the CMS

A correct ACL in the CMS is not enough: Graph evaluates access against the state available to it. That state can lag behind the CMS, and every cache between Graph and the browser adds a window of its own.

Series: Part 1: The ACL Is the Gate · Part 2: Keeping Graph in Step with the CMS · Part 3: Pages, Languages, Visual Builder, and What Broke (13 October)

Status: built and verified in September 2026 on CMS 13.1.2 and Optimizely.Graph.Cms 13.1.2 against a live Graph tenant, with .NET 10, Next.js 16 and @auth0/nextjs-auth0 4.30. Verified again on 13.2.0 of both, released on 21 September: the database schema upgrades on start, every test passes, and the findings that depend on the add-on version held, with one reservation: the reindex tail described just below. The preview measurements and the changes that came out of them are from 1 and 2 October, on 13.2.0; the ledger’s own secret and its behaviour without the CMS, from 3 and 5 October. 13.3.0 of both, released on 5 October, is not tested here. Vendor claims link to vendor docs; everything else is either measured or a design choice, and says which. A CMS 12 note is in part 3.

Three times on 13.2.0, an access change I had saved in the CMS took minutes to reach Graph’s answers. Twice, right after a full verification run, a teaser I had closed to Everyone with a plain ACL save stayed readable with the single key for more than two minutes. Once, on 3 October, a teaser standing in for a section stayed readable for 254 seconds to the role I had just removed, in an ordinary series of 19 such narrowings whose other 18 took one to two seconds; nine more, made during a burst of about fifteen publishes, took two seconds or less. For access control the tail is what counts: a gate that closed within two seconds eighteen times out of nineteen still stayed open for minutes the nineteenth time. A section narrowed the same way stays readable for as long to the role that has just lost it. The CMS was right the whole time. What Graph returned had not caught up yet, and nothing in my setup told me so. The docs allow more than the usual few seconds for bulk updates, and for a change of access rights alone, which is not among the documented sync events (part 3), they give no latency figure; the cause here is not established. These times are measured end to end, from the CMS save to Graph’s answer, so they implicate no single component, and three occurrences are not a rate.

In short. Problem: the CMS decides at once who may read a section, but Graph enforces that decision as it stands in its index, and Graph’s answers can lag: [measured] usually by seconds, three times by minutes on my tenant. Consequence: for that long Graph can still return a section to a role that has just lost it, and every cache on the way to the browser adds a window of its own. What I built: [my implementation] the BFF opens a section only when Graph and a list of readers it fetches from the CMS every five seconds agree, and a job in the CMS compares Graph’s answers with that list every ten minutes. Cost: a dependency on the CMS behind a five-second poll per BFF instance, a scheduled job, and the rules and windows this part lists. Limits: the ledger covers only what goes through the BFF, and only while the BFF can reach the CMS; a caller around the BFF, or a section narrowed from Everyone while the CMS is away, waits for Graph.

This is part 2 of 3. Part 1 built the CMS side: every gate is a public teaser and a section with its own ACL, the ACL narrows before a publish and never widens, and a repair under a database lock restores it whenever anything else writes it. That makes the CMS right. Graph is another matter, because it enforces the ACL as it stands in its index, updated when the add-on reindexes the section. [documented] The HMAC page describes the check per item: an item is returned only if it includes a read permission for the user or for one of the roles (u:{USERNAME}:Read, r:{ROLE}:Read).

This part follows the indexed ACL to the browser: holding Graph to the CMS through the BFF, checking the index itself for every audience, signing and caching in the BFF, and the front end’s cache for anonymous visitors. Any of them can keep a section readable after its gate has closed, and none of them reports it.

Three terms from part 1 carry this part. The single key is Graph’s read-only key for public content: it returns published content that everyone may read. The HMAC key is a key and secret for server-side calls. [documented] Optimizely lists HMAC under admin access, “full access to all Graph resources without restriction” (authentication); the same kind of credentials manages webhooks (manage webhooks) and signs the sync API that writes content to the index and purges it (sync content data). A query signed with it reads everything, unless the cg-username and cg-roles headers narrow it to what that user or those roles may read. The headers only narrow that one query; the key itself keeps its full access. And every gate is a public teaser plus a section, a shared block with its own ACL, which Graph returns only to the roles in that ACL. As in part 1, [documented] rests on linked vendor docs, [measured] on my own setup, and [my implementation] on a design choice.

The whole path, with the windows numbered as in the table of windows below:

Figure 1: The whole path from the CMS to the browser Figure 1. The whole path from the CMS to the browser: two routes out of the CMS, one gate in the BFF where they must agree, and the windows numbered as in the table below.

One narrowing, step by step, for a role an editor has just taken off a section (the numbers are the ones this part measures or sets):

  1. T0. The editor saves the narrower audience. The CMS is right at once.
  2. T0 + up to 2 s. The CMS’s list names the new readers.
  3. T0 + up to 7 s. The BFF’s next poll fetches the list; from then on the BFF refuses the section to the dropped role, whatever Graph says. [measured] 1 to 6 s to the front end in the browser test.
  4. Seconds, three times minutes. Graph’s answers catch up. Until then the single key, or a signed caller around the BFF, still gets the old answer.
  5. Up to 30 s after step 3. A page the front cached for anonymous visitors before step 3 can be served for the rest of its lifetime; when the changed list has the fronts revalidate, it goes within seconds of step 3.
  6. Within ten minutes and 30 s, plus the run. If Graph’s answers still disagree with the CMS, the reconciliation notices and asks for a reindex.

A role taken from a user, rather than from a section, holds until the token is refreshed (window 4).

When the index is behind: an audience ledger

Graph’s answer is right only once the change has reached it, and on 13.2.0 an ACL-only change usually reached Graph’s answers within a few seconds, but three times it took minutes. For that long Graph would still open a narrowed section to the role that has just lost it.

So the BFF checks Graph against a snapshot of the CMS’s own access rights. The CMS stays the source of truth. The CMS publishes who may read each live gated section at GET /gating/audiences: readers straight from tblContentAccess, the CMS’s internal table of access entries, and the sections from tblContent, built at most every two seconds, with an ETag. [documented] Optimizely’s docs say not to access CMS tables directly and do not guarantee them across versions (install database schema). [my implementation] I read them as a deliberate exception to that advice, for the reason in part 1: a cached read can be behind the database exactly when it matters. The ledger is a compensating control of my own, added after I observed the lag, not a pattern Optimizely documents. The BFF fetches the list every five seconds and signs its request with a secret the two share for this list alone (Gating:AudienceLedgerSecret), which may not equal the file ticket secret from part 1; an unsigned request gets a 404. A section Graph returns is opened only if the list also names one of the caller’s roles, the same roles the BFF sends as cg-roles (Everyone for an anonymous caller, who is served with the single key and sends none). A section a current list does not name is refused. A BFF that has not loaded a list yet holds back every section Graph opened. The visitor sees “temporarily unavailable” rather than a refusal, and the BFF reports itself not ready. The filter runs after the BFF’s cache, so a cached answer is held to the current rights too.

The list names only what a visitor can carry: the roles from the role list (see Roles, below), plus Everyone and Authenticated. A section’s ACL also grants editorial roles, the managed editors from part 1 and, with the option part 3 describes, an approval group an administrator may add. Those let editors work on the section in the CMS. The BFF never sends them for a visitor, so the list leaves them out. That also covers a mismatch I would rather not rely on never happening: if the BFF got a role list with a new name in it before the CMS did, and that name was already an editors’ group, the list would still not let a visitor in through it.

The ledger cannot open anything: a section needs both Graph and the list. A stale list errs both ways, within limits: it delays grants and revocations alike. A section published since the last fetch is refused, and a role dropped since then keeps the section until the next refresh that succeeds, two minutes at most, and only while Graph still returns it too. The exception is Everyone: a stale list still serves the sections it showed open to Everyone (below), so a section narrowed from Everyone while the CMS is away stays open past the two minutes, for as long as Graph returns it. The ledger needs Gating:FileBaseUrl and its secret. Outside development a BFF without them does not start, unless the ledger is switched off on purpose (Gating:AudienceLedger=false); it used to switch itself off and only warn. It also costs a dependency. A BFF that loses the CMS keeps its last list for two minutes, refuses sections published since and sees no new narrowing. After that it stops trusting the list for restricted sections: it serves only the sections the list opens to Everyone and withholds the rest. A visitor the last list admitted is told “temporarily unavailable”, as in a Graph outage, not refused: without a current list the BFF cannot tell whether the caller may read the section, and the verdict, the metric and the audit record say so. The BFF’s readiness check then reports Degraded, not Unhealthy, for the reason a Graph outage does (see Operations below): every instance shares the CMS, and a check that failed on all of them at once would let a load balancer take the whole site down, public pages included. Only an instance that has had no list since it started reports Unhealthy, so it takes no traffic until it has one; one that starts while the CMS is down opens no gate. When the list changes, the BFF has the front ends revalidate (Gating:FrontRevalidateUrls), so an anonymous page cached before a narrowing is not served out its lifetime.

In the recorded runs of the browser test, an editor narrows a section and the dropped role loses it on the front end 1 to 6 seconds later; the test allows 20, and before the ledger it waited up to 150 for the index and the BFF’s cache. Each time the ledger overrules Graph it counts it (gating.ledger.withheld, by reason); the reason narrower-in-cms counts the views in which an answer from Graph, fresh or from the BFF’s cache, opened a section the CMS had already narrowed; in a production deployment it would be the first measure of that lag, and an upper bound on the index’s share of it.

It does not cover the path around the BFF. The single key reads only content everyone may read (documented). [measured] On the tenant it returned the one section open to Everyone and that section’s copy indexed for a fallback language (part 3), and none of the restricted sections the unrestricted HMAC key returned. That leaves the sections the index last saw readable by everyone: a section open to Everyone narrowed to a role, or one opened for a moment by a save to subitems, if the reindex after the repair lags. In my runs, such a section stayed readable with the single key until Graph’s answers caught up. Here the single key lives in the BFF’s and the CMS’s configuration and never reaches a browser; where it is published, narrow a public section by creating a new one, or rotate the key.

[documented] Graph also caches answers. A query that asks with items loses its cached answer when any content is published; one that asks with item, only when that item is published again (common mistakes). The docs do not cover whether a reindex after a change of access rights alone clears either; [my implementation] so I treat a client that asks with item as one that may see a narrowed section for longer. My BFF asks with items.

Holding the index to the CMS

The ledger protects the path through the BFF. Nothing so far told anyone that the index itself disagreed with the CMS. And the index can also keep a document the CMS no longer has. [documented] A full synchronization only adds and updates the content it finds and does not delete what it failed to reach; a smooth rebuild, which provisions a new slot, removes such documents (what Graph indexes). [documented] A delete that reaches Graph marks the document _deleted, and only a signed request that asks for deleted content with cg-include-deleted: true gets it back (modified and deleted content items). [measured] My test index held two documents whose content had been deleted in the CMS, and Graph still returned them. The reconciliation does not need to know why.

So the CMS checks the index on a schedule. Every ten minutes a job asks Graph which gated sections it returns to each audience a caller can have (an audience here is a set of visitor roles, not a CMS Audience). The audiences are anonymous with the single key, signed in with no role from the role list (cg-roles containing only Everyone and Authenticated), and each role in the role list on its own, with those two added (see Roles, below); the last two go over HMAC with both headers the BFF sends, so Graph filters them by role: 24 queries here, one per audience while each audience sees at most 100 section versions. In general it is one query per 100 versions an audience can see, languages and fallback copies included, and twice that when the first comparison finds anything. [documented] Paging with skip reaches at most 10,000 hits (skip and limit); a larger site needs a cursor. [my implementation] Until then, my job fails any run that finds more than 10,000 for one audience, rather than compare part of the index. The job compares each answer with the list the ledger gets. A section Graph returns to an audience the CMS does not let read it is a finding, but only if a second full comparison after a 30-second grace still has it, so an editor’s change in flight is not drift.

What happens next depends on the kind of finding:

  • A live section the CMS has narrowed and the index has not (acl-behind) is repaired. The job saves its ACL again, read from the database under the section’s lock, which, as measured in part 1, has the add-on reindex it. The lock matters: a copy read before a narrowing and written after it would widen the section again.
  • A key with no live section behind it (stale-document) is only reported. Removing it takes a smooth rebuild of the index, which is an operator’s call. [documented] Until the new slot is accepted or abandoned, content updates go to it and are not visible on the live site (smooth rebuild), so a narrowing made during a rebuild, and the job’s repair of it, reach queries around the BFF only after Accept; the ledger still holds the path through the BFF.
  • A live section with no ACL of its own (inherited-acl) is only reported; the ACL repair from part 1 owns it.
  • Sections the CMS opens and Graph does not are only counted: usually a widening not indexed yet, the safe direction. A section that stays narrower run after run is one the index never got, and its readers are refused until it is published again.

Every finding is counted (gating.index.wider, by reason), and so is every run; an alert fires when no run has reached the comparison for 30 minutes.

[documented] The docs also describe two explicit ways to have content reindexed: the IContentIndexer service, from code (synchronize content events), and on CMS 13 the “Synchronize with Graph” job, an immediate, direct synchronization (scheduled jobs). [my implementation] The repair still saves the ACL again, the path part 1 measured; I have not tried IContentIndexer in the job.

To see it work on the live tenant, a script narrows a test section straight in the database, a write the docs rule out (above), used here only to simulate drift. That raises no event, so Graph keeps the wider ACL. In two runs:

  • in the report: the dropped role, and only it;
  • the job saved the ACL again;
  • Graph agreed with the CMS ten seconds after the run started, with the grace shortened to five seconds for the test and the run started by hand, outside the schedule;
  • the database ACL stayed narrow.

On the demo’s normal state the report is clean: 24 audiences, zero wider, zero narrower.

For queries around the BFF, the reconciliation bounds how long an index that disagrees with the CMS goes unnoticed: ten minutes, plus the 30-second grace and the run itself. Then it asks for a reindex. It cannot make Graph take that reindex any faster, and during a smooth rebuild it cannot reach the live index at all, so the window stays open until Graph has caught up. Until then, a section the index last saw readable by everyone is still readable with the single key. The job runs in the CMS, which holds the HMAC keys and reads access rights from its own database. The CMS 12 version of my code does not have it (part 3 has the CMS 12 note).

Put together, these are the windows in which a closed gate can still open:

#PathHow long a closed gate can stay openWhat bounds it
1a section’s text through the BFFuntil the BFF’s next list: about two seconds in the CMS plus a five-second poll ([measured] 1 to 6 s in the browser test); while the CMS does not answer, up to two minutes, after which every restricted section is withheld; a section narrowed from Everyone meanwhile, until Graph’s answers catch up (row 5)the ledger; for a section narrowed from Everyone while the BFF cannot reach the CMS, nothing but Graph itself: the reconciliation runs in the CMS, so it helps only while the CMS is up
2an anonymous page in the front’s cacherow 1 plus up to 30 s, for a page cached just before the BFF’s list changed; row 1 plus seconds when the revalidation reaches that front instancethe lifetime, the revalidation
3a file link already issuedup to the ticket’s five minutes (part 1)the ticket’s expiry
4a role taken from a userup to the access token’s lifetime (5 to 15 minutes, below)the token refresh
5the single key, or a signed caller around the BFFuntil Graph’s answers catch up: seconds as a rule, minutes three times ([measured])nothing on that path; the index reconciliation notices it within ten minutes plus 30 s and the run, and asks for a reindex
6a draft, with Gating:AllowDraftIndexing, read by a signed caller without the status filteras long as the draft existsa rule only
7a document the index kept after its section was deleted, around the BFFuntil a smooth rebuildan operator

What the index holds, and what preview pays for it

Part 1 set ContentVersionSynchronizationMode to PublishedOnly and stopped the CMS on anything else. That keeps drafts of gated sections out of the index, and every other draft with them. Part 1 did not say what that costs: preview, for a site that renders it from Graph.

[documented] On CMS 13, live preview reads unpublished content from Graph: the front end sends the preview token from the preview URL as a bearer token, valid for five minutes (enable live preview), and the token is validated against the editor’s permissions (applications). Whatever credential the preview request carries, it cannot return a draft the index does not have. Optimizely’s guide to live preview with Next.js, written for “CMS 12 or later”, turns draft sync on with AllowSyncDraftContent, which I did not find in the CMS 13 docs; there the documented switch is ContentVersionSynchronizationMode, whose documented values include DraftAndPublishedOnly. [measured] On CMS 13 I varied only that switch.

[measured] On 13.2.0 with PublishedOnly, I saved a draft of a page and of a gated section, each with a marker in its name, and after thirty seconds asked Graph over HMAC without role headers, which per the HMAC docs returns content whatever its status. Both published versions were there, but no version carried the marker, so a front end that renders preview from Graph has nothing to render.

My front end has no preview route, which is why my tests never noticed. A real site has editors who expect one.

[my implementation] So drafts are now an explicit opt-in. With Gating:AllowDraftIndexing=true, both startup guards from part 1 also accept DraftAndPublishedOnly; All still stops the CMS, because it adds previously published versions as well (part 1). [measured] With the opt-in, on the same tenant and with the same marker:

Who asksThe section’s draft
HMAC without role headers, any statusin the index, with the marker
the single keynot returned; the section is gated
a query shaped like the BFF’s: role headers and status: Publishednot returned; only the published version

What a signed query with role headers and no status filter returns is the Graph service’s own behaviour, which this series leaves out; the rule below is a precaution based on the documented contract, and it treats that case as unsafe.

The price comes as a rule, with no mechanism behind it. [documented] The HMAC docs say a request without role headers returns content whatever its status, and they do not say that role headers change that. So treat a signed query without the status filter as one that can return a section’s draft to the section’s audience before anyone publishes it. Every query my BFF sends filters on status: Published, for pages and for sections. Anything else you let query the tenant with the HMAC key has to do the same. [documented] The preview route itself needs no HMAC key: it reads drafts with the editor’s preview token (above). If you cannot promise that for every caller, keep PublishedOnly and render preview in the CMS itself, without Graph.

The BFF: signing, headers, cache, resilience

A signature on every attempt. Requests with roles are signed with epi-hmac. Why not Basic, which the Graph C# SDK for CMS 13 uses: Basic sends the secret itself with every request, while an HMAC request carries only a signature bound to a timestamp and a nonce; [documented] Optimizely recommends HMAC or JWT for secured production environments (authentication). If you sign by hand, take the recipe from Optimizely: the HMAC page gives the header format, epi-hmac APP_KEY:TIMESTAMP:NONCE:SIGNATURE, and the Sync content data page gives the signing itself, as a Postman script: apiKey + METHOD + path-and-query + unix milliseconds + nonce + base64(MD5(body)), HMAC-SHA256 with the base64-decoded secret. [measured] My BFF signs the same way, and my tenant accepts it for queries too. The signature carries a timestamp and a nonce, so every attempt should be signed afresh, and registration order is what makes that true:

client.AddStandardResilienceHandler(o =>
{
    o.AttemptTimeout.Timeout = attemptTimeout;                               // 2.5 s
    o.AttemptTimeout.TimeoutGenerator = _ =>                                 // 10 s until Graph has answered once
        ValueTask.FromResult(warmth.Warm ? attemptTimeout : cold);
    o.TotalRequestTimeout.Timeout = cold * 2 + TimeSpan.FromSeconds(1);
    o.Retry.MaxRetryAttempts = 1;
    o.Retry.Delay = TimeSpan.FromMilliseconds(200);
    o.Retry.ShouldRetryAfterHeader = false;

    // 429 is neither retried nor counted by the breaker (see below).
    o.Retry.ShouldHandle = args => ValueTask.FromResult(IsTransientFailure(args.Outcome));
    o.CircuitBreaker.ShouldHandle = args => ValueTask.FromResult(IsTransientFailure(args.Outcome));

    // A per-instance budget of Graph calls: a burst of cache misses gets 503s.
    o.RateLimiter.RateLimiter = args => budget.AcquireAsync(1, args.Context.CancellationToken);

    // The default waits for 100 calls in the window, which a quiet site never makes.
    o.CircuitBreaker.MinimumThroughput = 10;
    o.CircuitBreaker.FailureRatio = 0.5;
});

// Added after the resilience handler, so it sits closer to the wire and signs every attempt.
client.AddHttpMessageHandler<GraphSigningHandler>();

static bool IsTransientFailure(Outcome<HttpResponseMessage> outcome) =>
    outcome.Result?.StatusCode != HttpStatusCode.TooManyRequests
    && HttpClientResiliencePredicates.IsTransient(outcome);

The reverse order compiles too, signs only the first attempt, and a retry replays the same nonce. Two tests check it: one builds the host’s pipeline and sees a fresh nonce on the retry, the other builds the reverse order and sees the nonce replayed.

Headers on a request with roles. cg-roles: the user’s roles, with the roles they imply, plus Everyone and Authenticated. cg-username: a fixed value that identifies nobody; the visitor’s email could match user-type ACL entries, and then the answer, and the cache key with it, would no longer depend on roles only. This is deliberate: the BFF admits visitors by role, never as a CMS user. cg-include-expired: false and cg-include-deleted: false. [documented] The HMAC page says a signed request leaves out deleted and expired content by default, but the Create queries page says a request with auth headers or Basic keys returns expired content unless it excludes it, and that expired content still shows as Published, so a status: Published filter does not exclude it. [my implementation] So the gate relies on neither default: every signed request sends both headers as false. I do not send the third switch the HMAC page lists, cg-include-hard-deleted, and rely on its default there.

Every other signed caller. The BFF need not be the only thing holding the HMAC key: a search service, an export or a second front end can query the same tenant. None of them goes through the ledger, so each has to keep the BFF’s rules on its own: both role headers on every request, status: Published in every query, and cg-include-expired: false and cg-include-deleted: false sent explicitly. The ledger and the BFF’s checks cover only the requests that pass through the BFF. Each of them also holds more than a query key (see the HMAC key above): whoever takes one over can do what the key allows, not only read what these rules let through. [documented] On DXP, a front end hosted by Optimizely gets the Graph app key and secret in its environment (Next.js ISR caching and Graph webhooks), so there the front end itself is one of these callers. Where Graph’s OIDC option fits your provider, a token limited to roles is the smaller thing to hand out; I have not tried it (part 3).

Roles. The role list is a contract: one file that the CMS (the editor’s picker) and the BFF both use. A role outside the list never reaches cg-roles, so a token cannot introduce one. An ACL has no hierarchy, so if GoldPartners should also see SilverPartners content, the role list says so ("implies": { "GoldPartners": ["SilverPartners"] }, followed transitively) and the BFF adds the implied roles; editors pick the lowest role that should see the content. The same map carries a rename: the new name implies the old one until every section that names the old role has been replaced by one that names the new role. It has to be a replacement, because editing will not do: published audiences only narrow, and the rule compares names.

Cache. [my implementation] This is the BFF’s own cache, not Graph’s. It keeps Graph’s answers under the query, the variables and the sorted cg-roles (anonymous for the single key): with a cg-username that names no one, two callers with the same roles send Graph the same request (see Headers). Two checks run after this cache, on every request: the business condition, against the caller’s own claim, and the audience ledger, so a cached answer is still held to the current list. The store is HybridCache in each instance’s memory, for 60 seconds by default, with a key prefix derived from the Graph endpoint, the app key, the CMS major version (13) and the content model. A shared second level belongs to one environment; the prefix guards against a mix-up but does not keep environments apart. Roles come from the token, so revoking one takes effect at the next refresh: keep the token’s lifetime at 5–15 minutes.

Resilience. 2.5 seconds per attempt and one retry after 200 ms, so about 5 seconds per call once the connection is warm. Until Graph has answered once after a start, an attempt gets 10 seconds, because a cold connection here took 2.2 to 2.6 seconds to set up (1.8 s in the warm-up measured after the fix; see the table); the BFF makes that first call at startup and reports not ready until Graph has answered it (part 3 tells how that was found). A page is two calls in series, three when the site replaces the page’s language (part 3), and the front end waits 15 seconds. When the sections query needs more than one page of results (100 items each, one per section and language, skip and limit; the BFF reads up to ten), the BFF has no request-level deadline, so a slow tenant on a large page can exceed the front’s budget; add one if your pages are that large.

A 429 is HTTP’s “too many requests” (RFC 6585). [documented] Since 5 October Optimizely documents the limits again: per tenant, per endpoint and HTTP method, over a 10-second window, by default 1,500 GraphQL queries per window (3,000 with an elevated allocation), reported on every response in X-RateLimit-Limit and X-RateLimit-Window, with Retry-After on a 429 and the advice to wait that long and retry, backing off if retries keep failing; answers from Graph’s cache do not count (rate limits). [my implementation] I don’t retry a 429: a wait of several seconds does not fit a 5-second budget, and “temporarily unavailable” now serves the visitor better than a slow page. The breaker ignores 429s, or a burst of them would cut Graph off for everyone. [measured] On 2 October, every answer I inspected from my tenant, with the single key and with HMAC, carried X-RateLimit-Limit: 2500 and X-RateLimit-Window: 10, which is not one of the levels the page lists; the page also mentions custom limits and calls that header the source of truth, so read yours and do not plan on mine. [documented] A 503, by the same page, is a heavy request that ran out of CPU time, not rate limiting; simpler queries or less concurrency help, and my one retry of a 5xx does not. Each instance has a per-second call budget, set explicitly outside development as this BFF’s share of the tenant’s limit, divided by its instances; anything else that queries the same tenant spends the same limit. The budget sits outside the retry, so it counts calls, not attempts: a retried call is two requests to Graph. It protects the quota, and rate limiting at the edge is still needed.

To the visitor, a transport failure is “temporarily unavailable” and never a refusal, whether Graph or the CMS’s list is what failed. A broken query stays an explicit 500, because that is a bug. A 401 or 403 from Graph, a revoked or rotated key, is counted as its own failure reason, so it reaches an alert as well as the log.

Operations. /healthz says the process is alive and never asks Graph. /readyz asks whether this instance can use Graph with its keys: a { __typename } query with the single key and one signed with HMAC. A 401 or 403 is Unhealthy; a 5xx or a timeout is only Degraded, because taking every instance out at once would turn a Graph blip into an outage; the ledger follows the same rule. A refused key is different: every instance shares the keys, so a revoked or rotated key takes all of them out of rotation at once. I accept that on purpose, so that a deployment with a wrong key never takes traffic, and it makes a key rotation an outage until every instance has the new key. Traces start at the BFF and follow each request to Graph, with Graph’s URL exported without its query string, where my BFF puts the single key (the single key page also documents an Authorization: epi-single header, which keeps the key out of URLs altogether); the BFF continues a traceparent if the front sends one, and my front does not yet. Every gate decision is written as an audit record: the gate, the section, the verdict, the trace id, and the subject as a hash only, keyed when Gating:AuditKey is set; set it outside development, because an unkeyed hash of an identifier can be reversed by guessing. The records go to their own log category, GatedContent.Audit, at Information in every environment; ship that category somewhere durable, because the BFF exports only metrics and traces.

Next.js: identity on the server, a cache for anonymous visitors

The identity provider. Auth0 is incidental. The BFF needs an OAuth 2.0 access token (a JWT) that carries the roles and email_verified as namespaced claims (Auth0:ClaimsNamespace, read with MapInboundClaims = false and a 30-second clock skew), put there by an Auth0 Action because an access token carries neither by default. Any provider that can add such claims plays the same part. For the BFF, another OpenID Connect provider needs only configuration (Identity:Oidc): the claim names, a map from the provider’s role values to the site’s role ids, and an explicit choice to trust a directory that issues no email_verified; without that choice, or a claim to read it from, the BFF does not start, and with a claim named that the directory never issues, every caller is anonymous: either way it fails closed. That path is tested with tokens signed by a key made for the test, not against a real Entra ID tenant. The front end still signs in through the Auth0 SDK; another provider needs its own sign-in code there. [documented] Graph can also accept a token from an OIDC provider directly, with the roles in a roles claim, once Optimizely Support has issued a separate key for the configuration (generic OIDC provider); I have not tried it (part 3).

The token. In @auth0/nextjs-auth0 4 the access token lives in the encrypted session cookie, but the SDK mounts /auth/access-token by default, which hands it to any script on the domain. I switch it off (enableAccessTokenEndpoint: false) and read the token in server code. In a production build, instrumentation.ts checks the configuration when the server starts: Auth0 complete (or persona mode asked for by name with ALLOW_DEMO_PERSONAS=true), an https BFF address (plain http only on loopback, since the token travels on it), and the revalidation secret. A missing variable stops startup, so nothing switches modes silently.

Who refreshes it. A Server Component cannot write cookies, so a token refreshed during a render cannot be saved, and with refresh-token rotation a refresh that is thrown away can end the session. The refresh therefore happens in proxy.ts, Next.js 16’s successor to middleware, on every request the SDK passes through: auth0.getAccessToken(request, response) refreshes an expired token and writes the new session onto the response, which the page rendering that request reads. The page only reads the token from the session; one still expired there means the refresh failed, and the visitor gets the public page and a sign-in prompt. A refresh token Auth0 has rejected ends the session at once. The login asks for offline_access, without which there is nothing to refresh with.

Gated pages are dynamic. A statically rendered page would land in the route cache, and one visitor would see content unlocked for another. Gated routes are force-dynamic; requests to the BFF use cache: 'no-store' and AbortSignal.timeout. This assumes cacheComponents is off; with it on, route segment dynamic is not available, and the same is done by reading the session at request time.

A cache for anonymous visitors. Every anonymous visitor gets the same answer, so it is kept in unstable_cache with a tag and a 30-second lifetime. With the default file-system store (measured here; see unstable_cache), a stale entry is served at once and refreshed in the background, entries on disk survive a restart, and tag invalidation lives in process memory. So every answer carries the time it was fetched, and code rejects one older than the lifetime:

let answer = await anonymousAnswer(url);            // unstable_cache with a tag; 503 and 404 are never stored
if (Date.now() - answer.fetchedAt > ttlSeconds * 1000) {
  answer = await askBffOnce(url);                   // stale: one fresh request per URL at a time
}

Answers for sections open to everyone carry part 1’s five-minute file tickets, so a 30-second lifetime leaves every cached ticket at least four and a half minutes.

In Next.js 16, unstable_cache is superseded by 'use cache'; I have not moved to it.

Invalidation. A Graph webhook, registered on the tenant with a custom X-Webhook-Secret header, reaches one BFF instance. The route exists only when that secret is configured, and compares it in constant time. The instance clears its cache, passes the webhook to the other instances in parallel, then calls the front ends’ /api/revalidate, which run revalidateTag(tag, { expire: 0 }). Peers and fronts get separate time limits, independent of the sender’s connection. Outside development the addresses must be https, and if no peer list is configured, the BFF shortens its cache to 30 seconds and says at startup how stale an answer can be.

Secrets. Four shared secrets, each of at least 32 characters and each for one purpose: the webhook secret, the file ticket secret, the ledger secret and the revalidation secret. The BFF refuses to start with a shorter webhook or ticket secret, and outside development with a shorter ledger secret or one equal to the ticket secret; the front refuses a shorter revalidation secret; and the CMS treats a shorter ticket or ledger secret, or a ledger secret equal to a ticket secret, as none. Each has a current and a previous value, and both are accepted during a rotation, so a new value rolls out without a window of refusals. Here each Graph key has one value, so rotating the Graph keys is a coordinated restart of the CMS and the BFF; until the BFF has the new keys it gets 401s and /readyz reports Unhealthy on every instance at once, so behind a load balancer that honours readiness no page is served until then, and without one pages answer 500.

Rich text is rebuilt from an allow-list of tags in a single linear pass with a length cap. Editors are trusted, so this is only defence in depth; who may read a section is decided by its ACL, not here.

Outages. An App Router page that has rendered cannot answer 503, so the outage page goes out as 200 with noindex, and force-dynamic keeps it out of every cache.

Where this bites: three silent failure modes

Three failures on the way from the index to the browser that raise no error: the site keeps working, and visitors read sections their roles do not allow. Five more, in the CMS, are in part 1.

  1. An HMAC request without role headers is unrestricted. [documented] The HMAC page says: “By default, querying with HMAC is equivalent to querying as a super user.” Role filtering takes cg-username and cg-roles together; the docs do not describe what one of them does alone, so treat it as no filter at all. Build the headers in one place and test it, including the case of no roles at all (a fixed cg-username that names no one will do; see Headers).

  2. revalidateTag(tag, 'max') keeps serving a gate you just closed. The stale page goes on being served while it refreshes in the background. Use { expire: 0 } on the webhook path.

  3. Graph’s answers lag the ACL. In my runs on 13.2.0, measured end to end, usually by seconds, three times by minutes, and nothing in my setup told me. Pass what Graph opens through an audience ledger: the CMS’s own list of readers, polled every few seconds. And compare the index with that list for every audience on a schedule, because the ledger does not guard queries that go around the BFF.

From prose to proof

First, what these numbers are: one machine, one tenant and one person, over a few days. They show that the mechanisms work and roughly how fast, but say nothing about capacity.

Environment: Windows 11 workstation; CMS 13.1.2 (13.2.0, tested from 26 September; the ledger and the reconciliation only on 13.2.0), BFF and Next.js 16 on the same machine; a real Optimizely Graph tenant indexed from that CMS; Auth0 for identity (the browser test signs in with demo personas); dates 23 September to 5 October 2026.

Does the ledger close the gap?

MeasurementResultn
a narrowing reaches the front end through the ledger (browser test)6, 4 and 1 s; the test allows 20 s, and before the ledger it waited up to 150 s3
a narrowing made straight in the database, found and reindexed by the reconciliationthe dropped role only; Graph agreed 10 s after the run started, 5 s of it the shortened grace2
the reconciliation on the demo’s normal state24 audiences, 0 wider, 0 narrower1

Each mechanism has a negative control. With the ledger’s filter switched off, six of its tests fail. With detection switched off in the reconciliation, four of its seven tests fail; without the second comparison after the grace, the test of a change in flight fails. Without the longer timeout for a cold connection, the cold-start test fails. The two ordering tests for the signing handler check both registration orders.

What does a page cost?

PathResultn
BFF, cache hitp50 8 ms, p95 18 ms30
BFF, cache miss on a warm process (two Graph calls in series)19–160 ms3 role sets
warm-up call at startup (connection to Graph set up)1.8 s1
first visitor query after a restart0.3 s1

[documented] Graph answers repeated queries from its response cache, and those answers do not count against the rate limit (rate limits). I did not control for it, so the low end of the miss range may be Graph’s cache, not Graph’s work.

A cache miss costs two Graph calls per role set (three in a replaced language): that is the number to put against the tenant’s rate limit. I timed the Next.js page only on next dev, so there is no front-end number here.

Threats to validity

  1. One machine. The front-to-BFF hop was loopback; the BFF-to-Graph hop was a real internet hop from one location. A BFF in another region, or behind more network, will see different numbers.
  2. One tenant, one region. Graph timings vary by region and load.
  3. Few browser runs. The narrowing through the ledger reached the front end in 6, 4 and 1 seconds in three recorded runs; the test enforces 20 s.
  4. Small n. Three role sets for the miss, one warm-up after a restart, two timed reconciliation runs.
  5. Not deployed to DXP. Several BFF instances were not run together: the fan-out of webhooks to peers is covered by tests but was not measured; no Graph webhook reached my machine, so registering one with a custom X-Webhook-Secret header rests on the webhook docs; front-end revalidation on a changed ledger list has not been run live; and I have not tried the shared Redis cache Optimizely documents for several Next.js instances on DXP.
  6. I wrote the tests myself. Negative controls, and a pass after every change in which I tried to break my own code, mitigate that; they do not remove it.
  7. The CMS, the add-on and Graph can change. The ACL-save reindex the reconciliation relies on, and the reads of tblContentAccess and tblContent behind the audience list, which Optimizely’s docs advise against, hold on 13.2.0, where the ledger and the reconciliation were built; 13.3.0, released on 5 October, is not tested here. Re-check them on every upgrade, and after Graph’s monthly service releases (releases and channels). Part 3 lists the other version-dependent findings.

Decisions and what they cost

DecisionAlternativeCostWhen to change it
hand-rolled HMAC in the BFFOptimizely’s Graph C# SDK; or Graph’s OIDC option, which takes a token with a roles claim directly (part 3)a handler of my own and two ordering tests; the Graph C# SDK for CMS 13 authenticates with the single key or Basic and sends per-user roles over Basicwhen the SDK signs with HMAC
a cache per role set, 60 s in the BFF (30 s without a peer list), 30 s in the frontno cachethrough the ledger a narrowing reaches signed-in visitors within seconds; an anonymous page can be served from the front’s cache for up to 30 s, unless the BFF knows the fronts’ revalidation addresses: then a changed list revalidates them; on DXP, where instances behind a load balancer each keep their own cache, Optimizely documents a shared Redis cache (Next.js ISR caching and Graph webhooks)when anonymous pages must change instantly
an audience ledger in the BFFGraph’s answer alonea poll of the CMS every 5 s per BFF instance; a BFF that cannot reach the CMS refuses sections published since; after two minutes it serves public sections only and reports the rest as temporarily unavailable; a section narrowed from Everyone meanwhile stays as open as Graph has itnever, while the index can lag the CMS
the index reconciled against the CMS every 10 minutestrusting the indexone Graph query per 100 section versions per audience (24 here), twice that when a finding needs a second comparison 30 s laternever, while the index can disagree with the CMS
a 429 neither retried nor counted by the breakerhonour Retry-After and retry, as Optimizely advisesa burst shows visitors “temporarily unavailable”when pages can wait several seconds
drafts kept out of the index unless Gating:AllowDraftIndexingdrafts always indexed, or Allno preview through Graph by default; with the opt-in, every signed query must filter on statuswhen every HMAC caller filters on status, or preview renders in the CMS
the ledger lists only roles a visitor can carryevery role the ACL lets readeditorial roles are invisible to the BFF, by designnever

Three checks you can run in ten minutes

My code is not public, so here is what to check on your own site. Each check rests on the documented contract and works without my code:

  1. Log the names, never the values, of the headers on every HMAC request. Both cg-username and cg-roles must be there, every time.
  2. Call /auth/access-token on your front end, and read your webhook handler. The first must be 404; the second must use revalidateTag(tag, { expire: 0 }).
  3. Narrow a test section’s audience with a plain ACL save, then ask Graph over HMAC with the dropped role’s headers every 2 seconds. Ask in the query shape your site uses (item or items, which Graph caches differently). The time until it stops coming back is your window, and something in your design has to close it.

What’s next

The ACL in the CMS is right the moment it is saved; the one in Graph’s index may still be the previous one. So this part holds every section the BFF serves to the CMS’s current list, checks the index itself on a schedule, and bounds how long each cache can go on serving a closed gate.

Most of the gating machinery in parts 1 and 2 comes from two choices of mine. The audience lives in a property of the section, which is versioned; the section’s ACL is reflected in Graph, but is not versioned with the section itself. The narrowing before a publish, the rule that an audience may only narrow, the section lock and the database reads all exist to keep those two in step. And I let the ACL in Graph’s index decide who reads a section, taking access away included, which assumes more than the docs describe: [documented] a change of access rights alone is not among the events that sync content to Graph (part 3). So the ledger and the reconciliation exist to hold the indexed ACL to the CMS. Both are my choices, not Optimizely’s. Some of the machinery comes from neither: the teaser and the section, the placement rule and the file tickets follow from Graph filtering whole documents and from the CMS serving files under their own access rights. Part 3 tries two ways out, native access rights, which drop the first choice and keep the second, and a section the CMS serves itself, which drops both, and says what each one costs.

Part 3, out on 13 October, takes the pattern onto the structures a CMS 13 site has: pages on several sites, fallback and replacement languages, variations, Visual Builder compositions and containers, and puts it next to the CMS features that also touch a section: approvals, Opti ID, Opal, Audiences, scheduling and projects. It also collects what broke beyond the CMS side, the full table of decisions and their cost, and a sequence that works.

Further reading

All links checked on 2026-10-05. Optimizely moved its documentation to docs.optimizely.com at the end of September; some older bookmarks redirect to the new page, others only to the home page. The HMAC signing recipe is now on the Sync content data page, the CMS 13 page on enabling live preview has content again, and on 5 October rate limits got a page of their own.

Optimizely Graph

CMS 13

DXP

.NET, SQL Server, Next.js


Stack: EPiServer.CMS 13.2.0 and Optimizely.Graph.Cms 13.2.0 (the code was built on 13.1.2), .NET 10, SQL Server, Next.js 16, @auth0/nextjs-auth0 4.30, OpenTelemetry with Azure Monitor. Optimizely, Auth0 and Next.js are trademarks of their respective owners; this article is not affiliated with them. My code is not public.

// tech stack
EPiServer.CMS 13.2.0Optimizely.Graph.Cms 13.2.0.NET 10SQL ServerNext.js 16@auth0/nextjs-auth0 4.30OpenTelemetryAzure Monitor