Onboarding Microsoft Azure

Grant Ringleader a custom least-privilege role to run workstation VMs in one of your Azure resource groups, keyless, with Terraform or an ARM template.

Your developers’ workstations can run as VMs in one of your own Azure resource groups, on your bill and inside your network. This page sets that up.

You run one Terraform module, or one script plus an ARM template, against a resource group you own. It creates an identity for Ringleader to use, a custom role narrower than built-in Contributor, and tells Azure to let Ringleader sign in as that identity without a client secret. You send a few values back to Ringleader, and from then on your developers create and delete workstations themselves. Ringleader never holds a secret for your tenant, and never gets rights outside that one resource group.

What this creates

Onboarding creates two things in your Entra tenant:

  • An app registration called ringleader-workstations, with its service principal. This is the identity Ringleader signs in as.
  • A federated credential on that app. This is how Entra confirms that a sign-in really comes from Ringleader and is meant for your organization, without a client secret.

It creates two or three more in your resource group, all visible in the Azure portal afterwards:

  • A custom role called Ringleader Workstation Operator. It lets the holder create, start, stop and delete VMs and their disks and network interfaces, and nothing else. The full list is below.
  • A role assignment giving that role to the app, on this resource group only.
  • Optionally, a network for the workstations: a virtual network, a subnet, a NAT gateway with its public IP, and a network security group. Both paths create it unless you turn it off (create_network = false or CREATE_NETWORK=false).

What Ringleader can and cannot do in your resource group

There is no secret to leak. No client secret is created or stored, so there is nothing to rotate and nothing to steal. Each time Ringleader acts, it presents a short-lived signed token and Entra checks it. Delete the federated credential and Ringleader loses access immediately.

Other Ringleader customers cannot reach your resource group. Every token names the organization it is for, and your federated credential accepts only tokens that name yours. Entra turns the rest away; that decision is made in your tenant, not by Ringleader.

Ringleader can only manage VMs, in this one resource group. Built-in Contributor can touch every kind of resource in a group; the custom role lists only what running workstations takes:

  • Microsoft.Compute/virtualMachines/*: read/write/delete and power operations
  • Microsoft.Compute/disks/*: the OS-disk lifecycle
  • Microsoft.Network/networkInterfaces/* and publicIPAddresses/*: the per-VM NIC and optional public IP
  • Microsoft.Network/virtualNetworks/subnets/{read,join/action}: place a NIC on the workstation subnet (never modify the vnet)
  • Microsoft.Resources/tags/write: keep VM tags up to date
  • read-only size / usage / resource-group introspection

The role can be assigned only within your resource group, and it is. Three features are on by default and add to it: runtime identities, egress control and artifact storage. Each section says what it adds and how to turn it off.

If you write your own role instead: keep the tags action

Azure updates a VM’s tags through a separate action from virtualMachines/write. Leave out Microsoft.Resources/tags/write and workstations start fine, and then every later tag update fails with a permission error, possibly weeks afterwards. Built-in Contributor happens to include it; a narrower role must list it.

Before you start

You will have received two values from Ringleader when your organization was set up. If you do not have them, ask us.

ValueWhat it isExample
issuer URLThe address Ringleader’s tokens come from. No trailing slash.https://oidc-app.ringleader.dev
organization idYour organization’s id: a UUID, never its name.0192f5bf-af83-7178-8d0a-f1c7aea06bde

You also need a resource group to put the workstations in. You create it first; the module and the script do not. We recommend a new, empty one used for nothing else: workstations are billed to it and everything Ringleader may do is limited to it, so a dedicated resource group keeps the bill and the risk in one place. You also need permission to create an app registration in your Entra tenant, for this setup only.

Register one subscription feature first

Register Microsoft.Compute/UseStandardSecurityType on the subscription. Ringleader creates VMs with securityType=Standard, because the Trusted Launch default does not allow the nested virtualization workstations rely on. Without this feature every VM create fails: everything else will look right, and no workstation will ever start.

az feature register --namespace Microsoft.Compute --name UseStandardSecurityType
# poll until it says Registered (this can take several minutes), then propagate it:
az feature show --namespace Microsoft.Compute --name UseStandardSecurityType \
  --query properties.state -o tsv
az provider register --namespace Microsoft.Compute

Also confirm your chosen region has capacity for the VM size you intend to use; without it, every create returns SkuNotAvailable.

Set it up

Two paths that do the same thing, so use whichever your team already runs. Both live in github.com/ringleader-dev/cloud-onboarding, and both finish by printing the values you send back to Ringleader (see what you send back).

PathWhat it doesUse when
Terraformeverything, end to endyou manage infrastructure as code
ARMa script creates the app and its credential with az, then deploys the role, the assignment and the network as templatesyou prefer az or the portal

The ARM path needs the script because an app registration and a federated credential live in Entra, not in a resource group, and an ARM template cannot create them.

ARM

git clone https://github.com/ringleader-dev/cloud-onboarding
cd cloud-onboarding/azure/arm

export RG=ringleader-workstations                      # the resource group you created
export ISSUER_URL='https://oidc-app.ringleader.dev'
export ORG_UID='0192f5bf-af83-7178-8d0a-f1c7aea06bde'
export SSH_SOURCE_CIDR=203.0.113.0/24                   # where your engineers connect from

az login
./deploy.sh

deploy.sh is safe to run again if something goes wrong partway, and prints the values to send back. SSH_SOURCE_CIDR adds an SSH rule of your own; see Reaching your workstations for who can connect without it.

Onboarding a SECOND resource group in the same tenant? Set the names.

deploy.sh takes two more optional variables, and both matter as soon as a second resource group in the same Entra tenant is onboarded (a second team, a second environment, prod alongside dev). Left at their defaults, one fails loudly and the other fails silently:

export APP_NAME='Ringleader Admin (team-b)'                  # default: ringleader-workstations
export ROLE_NAME='Ringleader Workstation Operator (team-b)'  # default: Ringleader Workstation Operator
  • ROLE_NAME. Azure requires a custom role’s display name to be unique across the whole tenant, even when the roles are scoped to different resource groups. Re-using the default for a second onboarding fails the deployment with RoleDefinitionWithSameNameExists.
  • APP_NAME. deploy.sh finds its app registration by display name and reuses the first match, which is what makes re-runs safe. Left at the default, a second onboarding silently adopts the first onboarding’s app and adds the new federated credential to it, so two boundaries end up sharing one identity. Name the app per team or per environment and the lookup stays honest.

One resource group in a tenant needs neither variable.

Terraform

cd cloud-onboarding/azure/terraform/examples/standalone
cp terraform.tfvars.example terraform.tfvars
# Edit terraform.tfvars: your subscription id, the resource group you created,
# the two values above, and uncomment ssh_source_ranges with the addresses
# your engineers connect from.
az login
terraform init && terraform apply
terraform output handoff                        # the values to send back

The portal, step by step

The same result without deploy.sh. The app registration and its federated credential live in Entra, so you create them in the Entra portal; only the role and its assignment are a template.

  1. Create the app. Microsoft Entra ID → App registrations → New registration. Give it a name, choose Accounts in this organizational directory only, and register. Copy the Application (client) ID, which is the app client id you hand back.

  2. Add the federated credential. On that app: Certificates & secrets → Federated credentials → Add credential → scenario Other issuer.

    FieldValue
    Issuer<issuer-url>/org/<org-id>
    Subject identifierorg:<org-id>
    Nameringleader-oidc
    Audienceapi://AzureADTokenExchange

    All three values must match exactly, with no trailing slash on the issuer.

  3. Find the service principal’s object id. Microsoft Entra ID → Enterprise applications → your app → Object ID. This is not the client id, and it is what the template wants.

  4. Deploy the role. Portal search → Deploy a custom template → Build your own template in the editor → load azure/arm/azuredeploy.json → Save. Choose your resource group (the template is resource-group scoped) and set principalId to the object id from step 3. The template defaults enableWorkstationIdentities to true, so set it to false unless you use runtime identities.

If you are onboarding a second resource group in this tenant, give roleName a distinct value here too, as in the note above.

What Entra checks

When Ringleader signs in as your app, Entra checks three things about the token it presents. The first two are built from the two values Ringleader gave you; the third is a fixed value Microsoft documents. The module and the script configure all three, so you never type them yourself unless you use the portal.

What Entra checksValueWhere it is configured
Who signed the token<issuer-url>/org/<org-id>the federated credential’s Issuer
Which organization it speaks fororg:<org-id>the federated credential’s Subject identifier
What the token may be used forapi://AzureADTokenExchangethe federated credential’s Audience

Entra compares the issuer character for character, so a stray trailing slash is enough to fail. If a workstation fails to create with AADSTS700211: No matching federated identity record found on its status, one of these three does not match. Compare them against what was created, starting with the trailing slash.

Reaching your workstations

rl shell, rl tmux, port-forwards and VS Code Web all connect to the workstation over SSH on TCP 22, straight to the machine. A workstation behind an Edge is reached through a port on the edge instance instead.

Ringleader writes the rule that lets you in. It adds one rule to the network security group (NSG) on the workstation subnet. The rule admits TCP 22 and 2222 from any address to the workstations Ringleader creates, and to no other VM, unless the CloudAccount lists sshSourceRanges. It takes a priority from 3000 to 3999, so keep that range free in the subnet’s NSG. Writing it needs the egress control grant, including the actions on application security groups that section lists.

The module and the script can also add a rule of your own on TCP 22 and 2222, for the addresses your engineers connect from:

ssh_source_ranges = ["203.0.113.0/24"]   # Terraform, in terraform.tfvars
SSH_SOURCE_CIDR=203.0.113.0/24 ./deploy.sh   # ARM

Your rule covers every VM on the subnet, not only Ringleader’s workstations. It applies alongside Ringleader’s rule, and keeps workstations reachable from those addresses when Ringleader cannot write its own. A workstation with no egress policy is the exception. The NSG Ringleader puts on its network interface admits SSH only from the addresses Ringleader’s own rule allows, so your rule does not widen it.

A workstation cannot fully confirm any of this for you. It sets itself up over its own outbound connection, so it reports Ready whether or not anything can reach it. When Ringleader could not write its rule, the workstation reports the SSHAdmissionMissing condition. Otherwise, open a shell to check.

A public IP by default, and the NAT gateway

On Azure a workstation gets a public IP unless you say otherwise. An existing workstation keeps the setting it was created with, so one created while the default was no public IP still has none. A VM with no public IP has no way to reach the internet on its own. The network the module creates therefore includes a NAT gateway, which lets such a workstation reach Ringleader and finish setting up.

The NAT gateway and its public IP bill by the hour whether or not anything uses them. The only way to avoid that is to bring your own subnet that already has outbound access (create_network = false or CREATE_NETWORK=false).

If you would rather none of your workstations had a public IP, set allowPublicAddresses: false on the CloudAccount. To set it for the workstations one CloudIdentity builds instead, the Ringleader object that says how workstations in your resource group are built, set this on it:

spec:
  defaultProviderConfig:
    azure:
      publicIp: false

Use defaultProviderConfig when developers may turn it back on for a workstation, and overrideProviderConfig when they may not.

Runtime identities (on by default)

By default a workstation has no Azure identity of its own. Nothing running on it, including an AI coding agent, can touch anything in your subscription. That is the safe default, and there is nothing to narrow.

Ringleader can instead give your developers their own Azure identity to work as: it creates a user-assigned managed identity per user, shared across that user’s workstations, and assigns roles to it, so one person’s workstations can read one storage account and nobody else’s can (see runtime identity).

This is the one default on this page worth a deliberate decision, because creating those identities and assigning them roles takes two things built-in Contributor cannot do: manage managed identities, and write role assignments. Both are limited to your one resource group, but they are still the power to hand out access inside it.

So: in a resource group dedicated to Ringleader workstations, leave it on. In one that holds anything else, turn it off when you apply:

enable_workstation_identities = false     # Terraform
WORKSTATION_IDENTITIES=0 ./deploy.sh      # ARM

With it off, a workstation that asks for its own identity fails with a 403 error rather than silently getting nothing, and nothing else on this page changes.

Egress control (on by default)

By default a workstation can connect to anything your network can reach. Ringleader can narrow that to a list you choose per workstation, so a workstation can reach GitHub and your package registry and nothing else (see restricting outbound connections), enforced by network security groups Ringleader creates and keeps in step with the workstation’s configuration.

It is on by default. It also lets Ringleader write the rule that admits SSH to your workstations, and the NSG on the network interface of each workstation with no egress policy (see Two NSGs, at two layers). With it turned off, your workstations are reachable only through a rule of your own. A workstation with no egress policy still has no limit on where it connects. Turn it off with enable_egress_control = false (Terraform) or EGRESS_CONTROL=0 (ARM).

It adds twenty-three actions to the custom role, still scoped to your one resource group. A resource group onboarded with an earlier version of the module or template lacks the five application security group actions: apply the current version again to add them.

ActionFor
Microsoft.Network/networkSecurityGroups/{read,write,delete}one NSG per distinct egress policy, and the one NSG shared by workstations with no policy
.../networkSecurityGroups/securityRules/{read,write,delete}keep a policy NSG’s rules in step with the workstation’s configuration, and write the SSH rule in your subnet’s NSG
Microsoft.Network/networkSecurityGroups/join/actionattach the NSG to a workstation’s network interface
Microsoft.Network/routeTables/* (with routes/* and join/action)route traffic to the DNS/HTTPS proxy, for a policy that names hostnames
Microsoft.Network/virtualNetworks/subnets/{write,delete}attach the route table to a subnet and detach it again, which in Azure is a write of the whole subnet. A route table attaches per subnet, so routing is per subnet rather than per workstation. Ringleader never creates or deletes a subnet
Microsoft.Resources/subscriptions/resourcegroups/resources/readlist the resource group’s contents, so Ringleader can find and remove a gateway VM and its public IP that outlived their gateway; without it you keep paying for both
Microsoft.Network/natGateways/readsee whether a subnet already has an outbound path, before giving a gateway VM a public address of its own
Microsoft.Network/applicationSecurityGroups/{read,write,delete}, .../joinIpConfiguration/action and .../joinNetworkSecurityRule/actioncreate the group Ringleader puts each workstation’s network interface in, which the SSH rule it writes in your subnet’s security group names

Each distinct policy becomes one NSG shared by every workstation on that policy, rather than one per workstation. Azure caps an NSG at 1,000 rules and will not raise it, so one per workstation would not scale to a fleet.

Routing by subnet asks two things of your network, and neither is a subnet per policy. One gateway serves many policies from one subnet, telling them apart by source address.

  • A routed subnet may hold only workstations the gateway serves, because the gateway refuses a source address it has no rule for. The gateway does not start routing a subnet while a machine it does not serve is in it. A machine created there later is routed anyway, and loses its outbound access. create_governed_subnet (Terraform) or CREATE_GOVERNED_SUBNET (ARM) reserves a subnet for this. Put the workstations that have an egress policy in the subnet it outputs, governed_subnet_id (Terraform) or governedSubnetId (ARM). The workstations subnet the module creates is never routed, because it is for every workstation in the virtual network, with a policy or without one.
  • The gateway is never in the subnet it routes. A route table replaces the default route for everything in its subnet, so a gateway in the subnet it routes would send its own traffic back into itself. Its VM goes in the subnet that create_gateway_subnet (Terraform) or CREATE_GATEWAY_SUBNET (ARM) reserves. Give that subnet’s id, gateway_subnet_id (Terraform) or gatewaySubnetId (ARM), to the Edge as its subnet.

Two NSGs, at two layers, and which one is yours

Every workstation has two network security groups, one on the subnet and one on its network interface. Azure applies both, and traffic must be allowed by both:

WhereGroupManaged byDecides
the subnetthe one the module createdyou, plus the one SSH rule Ringleader addswho may reach the workstation
the network interface, with a policyone per egress policyRingleaderwhere the workstation may connect
the network interface, without a policyone shared by those workstationsRingleaderinbound from outside the virtual network: TCP 22 and 2222 only, from the addresses Ringleader’s SSH rule allows

Three things follow, and all three are about what you put where:

  • Keep your inbound rules on the subnet group. Apart from its one SSH rule, Ringleader leaves that group alone. On a workstation with a policy, the interface group admits all inbound traffic, so the subnet group alone decides who gets in. On a workstation without one, the interface group admits only TCP 22 and 2222 from outside the virtual network, even where your subnet group opens another port. So adding a policy to a workstation can let in whatever else your subnet group allows.
  • Do not put your own group on a workstation’s network interface. An interface holds only one, and if you bring your own interface (providerConfig.azure.networkInterfaceId) with inbound rules on its group, Ringleader replaces that group when it creates the workstation, policy or not, and those rules disappear. Put them on the subnet group instead.
  • Do not add an outbound Deny to the subnet group. It cannot make a policy tighter, because the interface group already blocks everything the policy does not list. It can only break a policy, by blocking something the policy allows, and that looks like Ringleader ignoring your list.

Artifact storage (on by default)

This grant lets Ringleader store files for your organization in a storage account in your resource group, instead of in Ringleader’s own storage, so the data stays in an account you control. Access is by Entra ID and a short-lived token, never a storage key.

It creates no storage account at apply time. Nothing uses this grant on Azure yet.

Two ways to take it, chosen by whether you name a storage account:

  • Leave artifact_storage_account_name empty and Ringleader creates and manages its own storage account and containers here.
  • Name a storage account you created and Ringleader gets blob access to it and nothing else, with no ability to create, change or delete a storage account. Its region, its lifecycle rules and its encryption key stay yours. Take this one if you have a data-residency or key-custody position to defend.

Turn it off with enable_artifact_storage = false (Terraform) or by setting the enableArtifactStorage template parameter to false (deploy.sh passes it through). With it off, those files stay in Ringleader’s own storage, and nothing else changes.

What you hand back to Ringleader

ValueHow to get itWhere it lands
app client idprinted by deploy.sh / terraform outputCloudAccount spec.azure.targetAppClientId
tenant idaz account show --query tenantId -o tsvspec.azure.tenant
subscription idaz account show --query id -o tsvCloudIdentity providerConfig.azure.subscriptionId
resource group, locationthe ones you usedproviderConfig.azure.resourceGroup, .location
subnet id (only if you created the network)printed by deploy.sh / terraform output handoffproviderConfig.azure.subnetId

Send them to whoever administers Ringleader for your organization, which may be you. They go into two Ringleader objects: a CloudAccount, which records which app to sign in as, and a CloudIdentity, which says how workstations in your resource group are built (subscription, resource group, location, VM size, subnet). See how Ringleader signs in to your cloud for a worked example of both.

Verifying

# The federated credential accepts only Ringleader's tokens, for your organization:
az ad app federated-credential list --id <client-id> \
  --query "[].{name:name, issuer:issuer, subject:subject, audiences:audiences}"

# The custom role has exactly the actions listed on this page:
az role definition list --name "Ringleader Workstation Operator" \
  --query "[0].permissions[0].actions" -o tsv

# The app holds that role on your resource group, and nothing else:
az role assignment list --assignee <client-id> \
  --resource-group ringleader-workstations \
  --query "[].roleDefinitionName" -o tsv

(If you set a custom ROLE_NAME, pass that name to az role definition list instead of the default above.)

Once an administrator has created the CloudAccount and CloudIdentity from the values you sent back, confirm Ringleader can use them before anyone creates a workstation. Replace azure and dev with the CloudIdentity’s name and namespace:

rl wait cloudidentity azure -n dev --for Ready --timeout 2m
rl cloudidentity get azure -n dev -o yaml     # status.valid: true

status.valid: true means Ringleader can use the identity. false means it cannot, and rl cloudidentity describe azure -n dev says why. The usual causes are a wrong app client id or a mismatch in one of the three values Entra checks.

Revoking

# Remove Ringleader's access but keep the app:
az ad app federated-credential delete --id <client-id> \
  --federated-credential-id ringleader-oidc

# Or remove the app entirely:
az ad app delete --id <client-id>

Terraform: terraform destroy.

There is no client-secret alternative

Azure onboarding was once also possible with a client secret you created and handed over. That path is retired: Ringleader stores no cloud credential, and a CloudIdentity naming one is refused at apply. The keyless setup on this page is the only one.