Strong listing sync with out SCIM


“Who’s your User?”

— Master Control Program, TRON (1982)

If consumer id is necessary to your software, you want a method to handle it. Usually this implies consumer accounts, and also you create them when somebody indicators up.

But permitting workers to enroll in no matter SaaS app strikes their fancy is a administration nightmare. So organizations like methods to regulate which customers exist (or don’t exist) in your app.

That controls who can check in, however how about what they’ll do?

For that you just want roles, or teams, which, just like the customers, come from the group’s id supplier – Entra, Google, Okta, and many others. Groups kind the premise of Firezone’s entry mannequin. They decide who can entry what.

The means of getting customers and teams into your app is known as listing sync, and it is surprisingly tough to do nicely. In this put up we’ll cowl what listing sync is precisely, the main customary for implementing it, and why we opted to forgo it totally to construct our personal engine from scratch.

An group’s listing is the checklist of who works there and what they’re allowed to entry. It lives within the id supplier and is fabricated from three issues:

  • Users: the folks within the group. Each has a reputation, an e-mail, and a standing like lively or suspended.
  • Groups: named collections of customers, like Engineering or Product. Groups allow you to grant entry to many individuals directly.
  • Members: who belongs to what. A member of a gaggle is both a consumer or one other group. A consumer’s entry comes from each group they belong to, straight or via different teams.

Directory sync is the method of copying that listing into your app and retaining it updated. The id supplier is the supply of reality, so what your app holds is a replica. Over time that replicate has to select up three sorts of adjustments: new customers and teams, updates like a reputation change, and removals. When somebody joins, leaves, or adjustments groups, ideally your app finds out rapidly.

It’s price clarifying what listing sync is just not: Single Sign-On (SSO). This could seem apparent however many people conflate the 2. Leading authentication requirements like OpenID Connect say nothing about easy methods to deliver customers into the appliance. In listing sync we’re referring to exactly the mechanism of bringing customers (and teams) in, however not how they’re authenticated (for that, see this post).

A simple example

Say your group has the next listing:

In your database it would look one thing like this:

The instance listing as database tablesThree tables. Users lists Alice and Bob. Groups lists Product and Engineering. Members has a consumer column, a gaggle column, and a guardian group column. Alice is a consumer in Product. Bob is a consumer in Engineering. The group Engineering is in Product. Each row leaves one of many first two columns empty, and answering whether or not Bob is in Product means following guardian teams via recursive queries.UserstitleAliceBobGroupstitleProductEngineeringMembersconsumergroupparent_groupAlice–ProductBob–Engineering–EngineeringProduct

To reply widespread entry questions in your app like “is Bob a member of Product?”, you’d have to have a look at Product’s members, then the members of any teams in there, and so forth till you discover Bob.

To keep away from that, we will flatten teams, giving every consumer a row for each group they belong to, straight or not:

The instance listing with flattened membershipsThree tables. Users lists Alice and Bob. Groups lists Product and Engineering. Members has a consumer column and a gaggle column, with one row for each group a consumer belongs to, straight or not directly: Alice is in Product, Bob is in Product, and Bob is in Engineering. Bob is in Product as a result of Engineering is nested in Product. Nested teams not seem as rows, so asking whether or not Bob is in Product is a single lookup.UserstitleAliceBobGroupstitleProductEngineeringMembersconsumergroupAliceProductBobProductBobEngineering

Now the identical query is a straightforward lookup.

Ok in order that’s the information mannequin. Now we have to get it from the id supplier into our database, and maintain it updated.

Of course, you do not have to invent all of this your self. There’s a regular for doing this listing sync factor. It’s referred to as SCIM and here is the way it describes itself:

The System for Cross-domain Identity Management (SCIM) specification is an HTTP-based protocol that makes managing identities in multi-domain situations simpler to help through a standardized service.

In apply this implies you construct an inventory of REST endpoints in your app, like:

GET /Users
GET /Users/{id}
POST /Users
PUT /Users/{id}
PATCH /Users/{id}
DELETE /Users/{id}

GET /Groups
GET /Groups/{id}
POST /Groups
PUT /Groups/{id}
PATCH /Groups/{id}
DELETE /Groups/{id}

And the id supplier calls them to maintain the listing in sync.

The promise of SCIM: a single API to sync a number of suppliers

The good factor about SCIM is that the id supplier pushes adjustments to your app, so the lag time between when an replace occurs on the supplier and when it lands in your app might be a lot decrease than a pull-based method that polls the supplier on a schedule.

Sounds nice on paper. But what are the issues right here?

Well, for starters, your app needs to be alive and able to obtain listing updates just about on a regular basis. Going down or getting overloaded for even a couple of a seconds means you may miss a crucial listing replace. Triggering a full sync out of your app is just not attainable, so that you lose these updates till the id supplier decides to inform you about them once more.

But the extra annoying factor about SCIM is, whereas it does an amazing job at standardizing the protocol (i.e. the wire format), it says nothing about how these endpoints ought to be referred to as, what lifecycle occasions they map to on the supplier, and even what sort of knowledge every request incorporates.

Identity suppliers all differ wildly in how they implement SCIM. Here are some examples:

  • Group membership updates: Okta can send additions and removals collectively in a single PATCH, whereas Entra requires them to be break up and solely permits one member removing per PATCH.
  • Even fundamental varieties differ throughout the similar supplier: Entra traditionally despatched lively: "False" as a string as a substitute of the SCIM-defined boolean, and solely later added a compatibility flag to change to compliant habits.
  • Optional options range: Okta explicitly doesn’t use a number of SCIM capabilities, together with bulk operations, POST searches, /ServiceProviderConfig, and filtering on meta.lastModified.

And in the case of consumer deprovisioning, the crucial operation of chopping off a departing worker’s entry, much more inconsistencies come up:

All of this boils right down to the truth that you find yourself with many various provider-specific SCIM code paths as a substitute of the one constant implementation you have been hoping to write down.

The actuality of SCIM: provider-specific code paths to clean over inconsistencies

It’s totally attainable you may find yourself with a extra sophisticated (and brittle) implementation going the SCIM route than in the event you constructed a pull-based system that hits every supplier’s API individually.

Which is precisely what we did.

With pull-based sync, you do the calling. Your servers hit the id supplier’s API on a schedule (or on the behest of the consumer), stroll the listing, and reconcile it along with your database. Every supplier’s API is somewhat completely different, however the form is at all times the identical, so many of the work might be shared.

How it works

Identity suppliers have API endpoints for itemizing customers, teams, and a gaggle’s members. After authenticating, a easy sync algorithm could possibly be:

  1. Get all teams
  2. Get all of the teams’ members
  3. For consumer members, get the consumer information similar to these members
  4. For group members, return to (2)
  5. Repeat each N minutes

A couple of particulars to bear in mind:

  • Pagination. Each checklist name returns one web page of outcomes, so “get all” means following the next-page hyperlink till there are not any extra.
  • Stable IDs. Every consumer and group has an ID that by no means adjustments. Match information on that as a substitute of e-mail or title, so a rename updates a document as a substitute of making a reproduction (see the Alice problem).
  • Cycles. A bunch can find yourself inside itself (A in B in A). Remember which teams you have already visited so step 4 would not loop eternally.
  • Flattening. This is the place the flattened desk comes from. When you discover Engineering inside Product, each member of Engineering additionally will get a row for Product. It additionally means a removing wants a recompute as a substitute of a delete, since Bob may nonetheless attain Product one other manner.

Simple sufficient. For tiny directories like our instance, you may get away with strolling the whole listing, gathering all the information up entrance, and writing it to your database in a single go.

Larger ones are somewhat trickier.

Syncing larger directories

Larger directories take extra time to stroll. And the longer the listing takes to stroll, the better the danger of shedding the accrued state earlier than you handle to flush it to the database. Deploy on the unsuitable time and you need to begin over, delaying the time till new listing knowledge makes it into your database.

What you are able to do to alleviate this considerably is to initialize every sync with an epoch, checkpointing the information as you go, then lastly eradicating all information older than the epoch:

  1. Initialize a begin timestamp.
  2. Fetch one web page and write to database with timestamp.
  3. Repeat for remaining pages.
  4. At the top, delete all information older than the beginning timestamp.
A checkpointed full sync: write every web page because it arrives, then clear up on the finish

Size is not the one factor that makes this difficult. A couple of different issues present up as directories develop:

  • Rate limits. The supplier’s API limits is likely to be shared with each different app the shopper makes use of. Hit the API too exhausting and you may decelerate their different instruments. Go slower than you would like, and again off when the supplier tells you to.
  • Bad responses. Sometimes an API returns a normal-looking 200 with a part of the listing lacking. If your sync treats “lacking” as “deleted”, it will possibly take away a whole lot of actual customers. Set a restrict, or circuit breaker, on how a lot a single sync is allowed to delete. This can stop nasty surprises.
  • Nested teams. Providers disagree on who does the flattening, and a few return stale outcomes if you ask them to.
  • Overlapping jobs. If two syncs for a similar listing run directly, they’ll overwrite one another’s adjustments. The chance this occurs grows with the scale of the listing, since a earlier sync may nonetheless be working when the following one kicks off. Make positive just one runs at a time (a job queue with uniqueness constraints or advisory locks works nicely), and {that a} crashed job would not block the following one eternally.
  • Temporary vs. everlasting errors. A timeout or a 503 ought to be retried. A revoked credential should not, as a result of retrying simply wastes your quota and hides the actual downside. Your sync ought to inform the 2 aside and let an admin know concerning the second type.

Ask for less

The much less you fetch, the quicker the sync. Most directories have loads in them that has nothing to do along with your app, like contractors, service accounts, and a whole lot of teams for mailing lists and workplace areas.

Where you may, let the supplier do the filtering. Ask just for the customers and teams assigned to your app as a substitute of every little thing. That means fewer pages to learn and fewer calls to make, which additionally eases the strain on the speed limits.

You may cease right here and stroll away with a reasonably sturdy listing sync engine. But in the event you use the listing for entry guidelines (like we do in Firezone), you are vulnerable to lagging necessary updates between the sync schedules.

Consistency properties

This course of could be very a lot finally constant: finally your database will mirror the group’s listing because it exists within the id supplier.

The above course of is a full sync – stroll the listing, seize each consumer, group, and member, then reconcile it along with your database – insert what’s lacking and take away what’s gone.

Depending on the supplier’s fee limits and the scale of the listing, this sync could possibly be as fast as a couple of seconds, or typically greater than an hour. In that point, it is fairly attainable the listing has modified: a brand new worker onboarded, one other left the corporate, another person went on parental depart.

Lagging a brand new worker onboarding may depart the worker unable to check in. Annoying, to make certain. But lagging an worker termination is dangerous – till you pull that listing change in, the worker nonetheless has entry. Are they going to steal firm secrets and techniques?

SCIM was supposed to assist right here, however we noticed how that goes. What else can we do to cut back the lag?

What about delta syncs?

One method to scale back the lag is to cease strolling the entire listing. Remember the place you left off, and subsequent time solely ask the supplier what modified. Some suppliers have a characteristic constructed for this (Entra calls it a delta question).

We checked out this intently and determined towards it. The primary motive is the teams we flattened earlier:

  • No transitive teams. A delta tells you a gaggle’s direct members modified. It would not inform you what meaning for the teams above it. If Bob is added to Engineering, you’d must work out by yourself that he is now additionally in each group that incorporates Engineering. Get that unsuitable and Bob finally ends up with entry he should not have, or with out entry he ought to.

There have been different issues too:

  • Deletes. With the epoch method, eliminated information handle themselves, as a result of something we did not see this time is gone. A delta feed has to report removals explicitly, and suppliers do it in a different way. Miss one and a former worker retains entry.
  • Expiring tokens. A delta solely works from a saved token, and tokens can expire (Entra’s final at most seven days) or get misplaced. When that occurs you want a full sync anyway, so now you have got two code paths to keep up as a substitute of 1.
  • Mistakes stick round. A full sync fixes its personal errors the following time it runs. A missed delta stays unsuitable till somebody notices.

To us, delta syncs resolve the unsuitable downside. The purpose is not to keep away from the total sync. It’s to make the total sync quick and dependable sufficient to belief, and resolve the lag downside via one other means.

Well, it seems id suppliers helpfully provide their personal set of real-time APIs you should use to get pushed-based listing updates – no SCIM required! (Google, Entra, Okta, JumpCloud). As a bonus, these APIs usually ship you updates quicker than they’d usually arrive over SCIM. Entra paperwork as much as a 40-minute lag time for its SCIM updates, for instance.

Firezone makes use of a hybrid method that mixes these APIs with an optimized full sync method from above, giving clients the perfect of each worlds: periodic reconciliation of the whole listing to make sure updates aren’t missed, and with real-time triggers that present near-real-time updates for crucial adjustments. Since the real-time triggers deal with pressing adjustments, the total sync can run a lot much less usually as a security internet, saving API quotas consequently.

A full sync and real-time updates take turns via one queue

The final piece is ensuring the 2 do not overwrite one another. A easy serialization queue in your job processing system handles that simply superb. It additionally helps to deal with every real-time notification as a sign to re-read the thing from the supplier, as a substitute of blindly trusting what the notification says.

Directory sync has a whole lot of transferring components when you strive it on an actual group. Large directories, flaky APIs, suppliers that disagree on what “deleted” means, and jobs that step on one another all must be dealt with. SCIM leaves most of that to your app, as soon as per supplier, and takes away your capacity to ask the supplier what’s true when one thing seems to be unsuitable.

Pulling from the supplier’s APIs retains you in command of when to look, what to belief, and what to do when one thing appears off. Add the suppliers’ real-time hooks on prime and also you get most of what SCIM promised, with far much less baggage.

There are extra edge circumstances we did not cowl right here. For instance, how do you sync two (or extra) directories containing a few of the similar customers into the identical account? That occurs when a corporation strikes to a different id supplier, or when two corporations merge. Firezone helps this, and the id aspect of it’s lined in the Alice problem.

Directory sync is out there on our Enterprise tier right this moment. If this is a vital characteristic in your group, book some time to get a better take a look at the way it works and take a look at it out.



Source link