RFC: Cloudflare Provider v5 Migration for
cloudflare_access
| Field | Value |
|---|---|
| Author | Miko Hadikusuma |
| Status | Draft — scoping only. No implementation proposed yet; see Open Questions |
| Created | 2026-09-21 |
| Context | cloudflare_access in terraform-modules is
built on resources deprecated in provider v4 and removed in v5 |
| Prompted by | The precedence collision that reddened the hb-infra production apply — infrahive#1431 |
| Related | infrahive#1417
(the sweep that triggered it), infrahive#1408
/ infrahive#1416
(the staging import), terraform-modules#140
(the interim v4 workaround),
infra/common-infra/business_unit_1/shared (in-house v5
pattern) |
Table of Contents
- Summary
- Why Now
- Current State
- What Changes in v5
- The Actual Blocker: the Provider Pin Is Root-Wide
- Migration Mechanics
- Design Decisions for the v5 Module
- Risks
- Proposed Phasing
- Open Questions
- Out of Scope
Summary
cloudflare_access wraps
cloudflare_access_application and
cloudflare_access_policy. Both are deprecated in provider
v4 and removed in v5. The last hb-infra production
apply emitted 176 deprecation warnings against them.
The headline finding of this scoping exercise:
The module cannot be migrated on its own. The provider version is pinned per root, not per module, and v5 renames most of the Cloudflare surface those roots use — DNS records, Argo tunnels, access groups. Bumping a root to v5 to fix
cloudflare_accessbreaks everything else Cloudflare in that root at the same time.
So this is not “migrate a module.” It is “migrate each root’s entire
Cloudflare surface, root by root,” and cloudflare_access is
one part of each of those. That reframing is the main thing this RFC
exists to establish before anyone starts.
The second finding: most roots are not on v4 yet, and v5 requires v4.52.5 first. So the near-term work is a v4 normalization (Phase 0), which is independently valuable and is where the real downtime risk — the v3→v4 tunnel rename — lives.
Why Now
Three reasons, in increasing order of importance:
- Deprecation noise. Every apply against a v4 root emits the warning. It is currently ignorable.
- The precedence collision. The module hardcodes
precedence = 1forthb_allow_emailandprecedence = 2forthb_vpn_bypass, with no override. Onadmin-crm.thb.nuandadmin-bea-crm.thb.nu, precedence 2 was already held by a hand-made policy (“Production BotDojo Service Auth”), so the create failed withpolicy precedences must be unique (12130)and took the production apply red. #1431 works around it by dropping the bypass from those two apps. v5 fixes this class of problem structurally — see What Changes in v5. - We already have a working v5 pattern in-house.
infra/common-infra/business_unit_1/sharedis pinned to~> 5.0and uses the new model (filemage.tf,qa_dashboard.tf,wireguard.tf). We are not designing from scratch.
Point 2 is worth dwelling on. The immediate reflex — add a
precedence variable to the module — is throwaway work: in
v5 precedence does not live on the policy at all, so the variable and
its plumbing would be rewritten at migration. That is the reason this
RFC exists instead of a two-line PR.
Current State
Consumers
Ten roots consume the module across twelve instantiations, six pinned
tags and two provider majors. Module HEAD is v0.1.58. The
distinction matters for planning, because the unit of migration is a
root, not an instantiation — two roots instantiate the module twice, at
different tags, and both copies have to move together.
The cloudflare_tunnel module sits beside it in every one
of these roots, on its own spread of tags. It matters as much as the
access module: every tunnel tag below v0.1.42 still
declares cloudflare_argo_tunnel and pins provider
3.9.1.
| Root | cloudflare_access |
cloudflare_tunnel |
Provider |
|---|---|---|---|
bees-infra/…/development |
v0.1.16 |
v0.1.16 |
~>3.9.1 |
bees-infra/…/non-production |
v0.1.16 |
v0.1.16 |
~>3.9.1 |
bees-infra/…/production |
v0.1.16 |
v0.1.16 |
~>3.9.1 |
ha-infra/…/non-production/cloudflare |
v0.1.53 |
v0.1.36 |
~>3.9.1 |
ha-infra/…/production/cloudflare |
v0.1.53 |
v0.1.36 |
~>3.9.1 |
hb-infra/…/development |
v0.1.42 |
v0.1.42 |
~> 4.0 |
hb-infra/…/non-production |
v0.1.53, v0.1.30 (eligibility) |
v0.1.36, v0.1.30 (eligibility) |
~>3.9.1 |
hb-infra/…/production |
v0.1.53, v0.1.42 (eligibility) |
v0.1.42, v0.1.42 (eligibility) |
~>4.0 |
pd-infra/…/non-production |
v0.1.37 |
v0.1.37 |
~>3.9.1 |
pd-infra/…/production |
v0.1.25 |
v0.1.25 |
~>3.9.1 |
Note that several roots instantiate the module more than once (a main
cloudflare-access plus a separate
eligibility-cloudflare-access), at different tags.
pd-infra/…/development is not listed: its
cloudflare-access and cloudflare-tunnel blocks
are commented out.
Scale
From the hb-infra production state, one root alone:
| Resource | Count |
|---|---|
cloudflare_access_application.apps |
87 |
cloudflare_access_policy.thb_allow_email |
87 |
cloudflare_access_policy.thb_vpn_bypass |
10 |
| Total | 184 |
hb-infra production is the largest root, but not by an order of magnitude. Expect the repo-wide total to land in the several hundreds.
Cloudflare Accounts
| Environment | Account ID |
|---|---|
| development | 945798da4af85bf579fa656c4731b23a (bees-, hb- and
pd-infra) |
| non-production | e6c92b47ed5bb4609a2dbe7bbbaa3c4f |
| production | e6c92b47ed5bb4609a2dbe7bbbaa3c4f |
The development row is not universal: both ha-infra
development tfvars set cloudflare_account_id to the shared
e6c92b47… account. ha-infra development does not consume
the module today, but “development is isolated from production” only
holds per root.
Non-production and production share a Cloudflare account, and each account is shared across several roots besides. Both facts are load-bearing for the v5 design, because v5 policies are account-scoped objects — see design decision 2.
What Changes in v5
Resource renames
| v4 | v5 |
|---|---|
cloudflare_access_application |
cloudflare_zero_trust_access_application |
cloudflare_access_policy |
cloudflare_zero_trust_access_policy |
The model change
This is the substantive part. In v4, a policy belongs to one
application and carries its own precedence. In v5, a policy
is a reusable, account-level object with no precedence,
and the application owns an ordered list of policy
references.
Today:
resource "cloudflare_access_policy" "thb_allow_email" {
for_each = { for app in local.subdomains_with_email_policy : app.domain_url => app }
application_id = cloudflare_access_application.apps[each.key].id
zone_id = each.value.cloudflare_zone
name = "Terraform Allow THB Email"
precedence = 1
decision = "allow"
include {
group = [var.cloudflare_group_id]
}
}
The v5 shape, as already used in
common-infra/shared/filemage.tf:
resource "cloudflare_zero_trust_access_policy" "thb_allow_email" {
account_id = var.cloudflare_account_id # account-scoped, reusable
name = "Terraform Allow THB Email"
decision = "allow"
include = [{ group = { id = var.cloudflare_group_id } }] # attribute, not block
}
resource "cloudflare_zero_trust_access_application" "apps" {
for_each = { for app in local.access_apps : app.domain_url => app }
zone_id = each.value.cloudflare_zone
name = each.value.access_app_name
domain = each.value.domain_url
session_duration = var.session_duration
policies = [
{ id = cloudflare_zero_trust_access_policy.thb_allow_email.id, precedence = 1 },
# bypass, external policies, … in order
]
}
Three consequences worth naming explicitly:
includestops being a block and becomes an attribute holding a list of objects. Everyincludein the module is rewritten, not just re-indented. Theipcase is a shape change, not a rename: v4 takes oneip = [list of CIDRs], v5 appears to want one object per entry. With 18 VPN CIDRs that is 18 include entries. Exact v5 shape foripneeds confirming against the provider schema — the in-house example only coversgroup.- Precedence moves to the attachment. The hardcoded
1/2disappear from the policy and reappear as an explicitprecedenceon each entry of the app’spolicieslist — ordinary data the module can compute per application. That does not by itself stop two entries on one application sharing a value — Cloudflare still requires precedences to be unique per application — but it makes the whole list visible to the module, so a collision can be rejected at plan time instead of failing the apply as #1431 did (see the worked example). - Policies become shareable. 87 identical per-app “Terraform Allow THB Email” policies collapse to one reusable policy attached 87 times. That is a large simplification, but the 87 old objects do not migrate into the new one and they do not disappear on their own — retiring them is an explicit step, described under Retiring the legacy policies.
This also explains the artefact that started all of this. “Production
BotDojo Service Auth” showed the same policy ID on both CRM
admin apps: it is already a reusable v5-style policy, attached twice.
The module’s legacy app-scoped policies and the dashboard’s reusable
ones are already coexisting in the same account, and that mismatch — not
the hardcoded 2 — is the real source of the friction.
The Actual Blocker: the Provider Pin Is Root-Wide
versions.tf in the module currently allows
>= 3.9.1, < 5.0. Moving it to >= 5.0
forces every consuming root to the v5 provider. But the provider is
shared by everything Cloudflare in that root, and v5 renames most of
it:
| v4 resource | Declared in consuming roots | Declared in the modules | v5 |
|---|---|---|---|
cloudflare_access_group |
15 | — | cloudflare_zero_trust_access_group |
cloudflare_access_policy |
12 | 2 (cloudflare_access) |
cloudflare_zero_trust_access_policy |
cloudflare_access_application |
9 | 1 (cloudflare_access) |
cloudflare_zero_trust_access_application |
cloudflare_record |
7 | 1 (cloudflare_tunnel) |
cloudflare_dns_record |
cloudflare_argo_tunnel |
0 | 1 (cloudflare_tunnel < v0.1.42) |
— (renamed to cloudflare_tunnel in v4) |
cloudflare_tunnel |
0 | 1 (cloudflare_tunnel ≥ v0.1.42) |
cloudflare_zero_trust_tunnel_cloudflared |
These are resource declarations on plan at
40eab6b8, counted in the ten consuming roots only (the
hb-infra sub-roots such as cloudflared/ and
hivebook-*/ already use the zero_trust_*
names). They size the HCL rewrite, not the state work:
for_each multiplies most of them, and the module
declarations multiply per instantiation. For the state side, the
hb-infra production instance counts under Scale are
the better calibration.
So there is no path where cloudflare_access
migrates alone. A root is on v4 or on v5. The
cloudflare_tunnel module has to move in lockstep, as does
every inline cloudflare_record.
And most roots are not on v4 yet. Cloudflare’s v5 migration guide
requires every root to be on provider v4.52.5 before
moving to v5. Eight of the ten consuming roots are still on
~>3.9.1, and their tunnel module is pinned below
v0.1.42, i.e. still on cloudflare_argo_tunnel.
The v3→v4 module work already exists — terraform-modules
v0.1.42 renamed cloudflare_argo_tunnel to
cloudflare_tunnel and moved both modules to v4, and
hb-infra production picked it up in 220fa69f — but the
other roots were held back because the rename plans as a
destroy/recreate of the tunnel, which would recreate the tunnels and
churn DNS. So the first step for each root is a v4 normalization:
provider v4.52.5, tunnel module v0.1.42 or later, and the
tunnel rename handled without a recreate (state rm +
import against the same tunnel ID). That is Phase 0.
This is the single most important thing for planning: the unit of migration is a root, and the work per root is the whole Cloudflare surface.
Migration Mechanics
moved blocks across resource types
Since Terraform 1.8, a moved block can change a
resource’s type, not just its address, when the
provider implements MoveResourceState. The repo is on
Terraform v1.16.1, and Cloudflare v5.19+ implements it
for everything in scope here — the v5 migration guide marks
access_application, access_policy
(account-level), access_group, record →
dns_record and tunnel →
zero_trust_tunnel_cloudflared as automatic.
cloudflare_argo_tunnel is not on that list, which is one
more reason Phase 0 retires it on v4 first.
moved {
from = cloudflare_access_application.apps
to = cloudflare_zero_trust_access_application.apps
}
Cloudflare’s tf-migrate rewrites the HCL (including
include blocks to attributes) and generates these
moved blocks. That removes most of the import-ID and
cf-terraforming questions, and with them the concern that
~184 imports in hb-infra production alone would have to be hand-written.
A state move also has no detach step. Pin the v5 target to
>= 5.19 — that is where the state
upgraders landed, and anything earlier puts us on the guide’s
stepping-stone upgrade path.
removed + import (the declarative pattern
already used for #1408 and #1416) stays the fallback for anything the
provider does not move automatically. Whether that set is empty for our
roots is what tf-migrate --dry-run in the Phase 1 spike should establish.
Ordering hazard
An Access application with no policies denies everyone. Any sequence that detaches the old policy before the new one is attached is a production outage on that hostname. The migration must attach-then-detach, not detach-then-attach, and that ordering needs to be proven in development before it is trusted anywhere else.
The default tf-migrate output is that
outage. The module’s policies are app-scoped, and those do not
move automatically: tf-migrate drops each app-scoped
cloudflare_access_policy and emits a removed
block, but it does not add policies to the
application. Applied as-is, v5 sends an empty policies
list, which detaches everything, and Cloudflare then deletes the
orphaned policies. The tool output must never be applied without the
step below.
Retiring the legacy policies
App-scoped policies have no independent lifecycle: once detached from their application, Cloudflare garbage-collects them. So there is no orphan sweep to do — the hazard is detaching them too early, not failing to delete them. The guide’s “Keeping existing policies attached” pattern handles that in two applies:
- Pin the existing attachments.
removed { lifecycle { destroy = false } }on the legacycloudflare_access_policyentries, and setpolicieson each application to the existing policy UUIDs. The plan should be 0 to add, 0 to change, 0 to destroy. - Swap to the reusable policy. Create the new
reusable
cloudflare_zero_trust_access_policyand switch the application’spoliciesentries to reference it, in the same apply. The old app-scoped policies are detached and garbage-collected by Cloudflare.
Step 1 needs the existing UUIDs per application, which is where generation (from state or the API) still earns its keep. The garbage-collection behaviour is per the guide and is worth confirming in the Phase 1 spike. If it turns out a cleanup is needed after all, it must delete only the policy IDs recorded from this root’s state, each verified to have no remaining attachment — never by name or account-wide, since the account is shared across roots.
Design Decisions for the v5 Module
These are the choices the rewrite has to make. Flagging them now so they get decided deliberately rather than falling out of whoever writes the code first.
1. One reusable policy, or one per app?
The v5 model invites collapsing 87 identical email policies into one. That is cleaner and cheaper. But a single shared object means a change to it lands on 87 apps at once, with no blast-radius control. Recommendation: collapse, because the policies genuinely are identical and divergence between them today would be a bug, not a feature.
2. One Cloudflare account, many roots and environments
Policies are account-scoped, and our accounts are shared along two independent axes. This is the sharpest hazard in the migration and must be settled before any code is written.
Axis one — environments. e6c92b47… is
shared between non-production and production. A naively named
account-level “Terraform Allow THB Email” would be the same
object for both, so a staging change would silently alter
production.
Axis two — roots. The same account is also shared
across stacks. Of the ten consuming roots, eight point at
e6c92b47… and two at 945798da…:
| Account | Roots consuming cloudflare_access |
|---|---|
e6c92b47… |
bees-infra non-prod + prod, ha-infra non-prod + prod, hb-infra non-prod + prod, pd-infra non-prod + prod (8) |
945798da… |
bees-infra dev, hb-infra dev (2) |
Each root has its own state and applies independently. So if every root declares its own reusable “Terraform Allow THB Email” for its environment, eight roots contend for one logical object in one account — creating duplicates, or fighting over a single policy with no owner. Note this is not hypothetical bookkeeping: design decision 1 collapses 87 per-app policies into one per module instance, which is exactly what turns a harmless per-app object into a contended shared one.
Axis three — instances within a root. hb-infra
non-production and production each instantiate the module twice
(cloudflare-access and
eligibility-cloudflare-access). A module that declares its
own account-level policy creates one per instance, so a root-derived
name alone would give two objects with the same name in one root.
Two ways to resolve it, and the RFC does not pick one:
- Single owner. One root owns each shared policy; the rest consume its ID. Cleanest object graph — genuinely one policy per account per environment — but it introduces cross-root coupling that does not exist today. Roots apply independently and nothing shares remote state, so the consuming roots would need the ID passed as a variable or looked up by a data source, and an ordering constraint appears where there was none.
- Instance-scoped identity. Each module instance
names its policy after its root and instance (e.g.
hb-infra production / eligibility — Allow THB Email), passed in as a required name-prefix input, so the objects are distinct by construction. No coupling, no ordering, no owner to assign. Costs a handful of near-identical policies per account — five or so in production — which is cosmetic next to the 87 we are already collapsing.
Recommendation: instance-scoped identity, on the grounds that independent roots are the property worth protecting and the duplication is trivial. A root that would rather have one policy shared by its two instances can still declare it once and pass the ID in, if the module accepts an optional policy ID in place of creating one. Either way this needs an explicit decision before the Phase 2 module rewrite, because it determines the module’s policy-naming contract and is expensive to reverse once roots have adopted it.
3. The application now owns its full policy list
In v4, Terraform managed only the policies it created; a hand-made
policy on a managed app was invisible to it. In v5,
policies is a complete ordered list on the application, so
anything hand-made becomes drift Terraform wants to
remove.
Concretely: after migration, BotDojo’s policy on the two CRM admin apps would be planned for detachment on the next apply unless the module is told about it. This is a regression in safety relative to today and needs an explicit answer — most likely a per-subdomain input carrying external policy IDs and their desired precedence, so hand-made attachments are declared rather than clobbered.
This is the sanctioned shape, not just our workaround: step 1 of the guide’s in-place pattern (policies left out of Terraform, referenced on the application by UUID) is exactly the BotDojo case, and the guide says that to keep policies owned outside Terraform you stop after step 1.
Worked example: BotDojo
BotDojo is the case that forced this RFC, and it is a good test of any proposed design because it is not hypothetical. What the API actually returns for it:
{
"id": "269c4ec3-d3f4-4cff-983a-5bed6871e4e0",
"name": "Production BotDojo Service Auth",
"decision": "non_identity",
"include": [{ "service_token": { "token_id": "5cb8a234-a50e-46a2-b64f-07e1be80ed84" } }],
"exclude": [], "require": [],
"session_duration": "24h",
"reusable": true,
"precedence": 2,
"created_at": "2025-06-06T14:08:22Z",
"updated_at": "2025-06-06T14:08:44Z"
}Four properties matter, and together they rule out doing anything about it before the migration:
"reusable": true. This is already a v5-shaped object — one account-level policy attached to bothadmin-crm.thb.nuandadmin-bea-crm.thb.nuby reference, which is why the same policy id appears under both. The v4cloudflare_access_policyresource is app-scoped (it takes anapplication_id) and has no notion of a reusable policy. Importing it into v4 risks Terraform reconciling it to app-scoped, detaching it from the second app and breaking BotDojo there on a green apply. Unverified, and deliberately so — the sensible way to find out is a throwaway app, not production.The pipeline cannot manage the service token.
GET /accounts/{account}/access/service_tokens/{id}with the Terraform credential (cloudflare_access_app_api_token,prj-p-secrets-3c7d) returns1010 auth.forbidden. So Terraform can reference5cb8a234-…but cannot own it, and acloudflare_access_service_tokenresource is not reachable without widening that credential’s scope. Note also thatclient_secreton a service token is returned only at creation and cannot be read back on import, so managing the token is undesirable even with the scope.Worth recording the diagnostic trap alongside it: under the same credential the list endpoint returns
success: truewith an empty result rather than an authorization error. That response is inconclusive, not evidence of absence — it is consistent with the list being authorization-filtered, and it briefly read as “the referenced token has been deleted, so the policy is dead.” The resource-specificauth.forbiddenis what establishes the scope limitation; the empty list establishes nothing. Confirming the token exists needs the Zero Trust dashboard or a credential with account-level service-token scope.Its contents are trivial and static. One include, nothing in
excludeorrequire, a default session duration, untouched since June 2025. There is no hidden complexity to preserve — the difficulty here is entirely about the resource model, not the policy.We own it. Ownership was an open question when this RFC was drafted and has since been confirmed, which removes the “not ours to touch” objection but not the two structural ones above.
So BotDojo is exactly the shape design decision 3 has to serve: an externally-created reusable policy that we own, whose position we want to declare and whose contents we do not want to manage. The v5 form is direct:
The policies list is a property of each application, so
it is assembled per application and the external
entries have to be keyed the same way. Sketching it against the module’s
existing for_each over local.access_apps:
locals {
app_policies = {
for app in local.access_apps : app.domain_url => concat(
contains(app.access_policies, "thb_email")
? [{ id = cloudflare_zero_trust_access_policy.thb_allow_email.id, precedence = 1 }] : [],
# only the subdomains that declare one — admin-crm.thb.nu and
# admin-bea-crm.thb.nu here, nothing for the other ~85 applications
[for p in lookup(var.external_policies, app.domain_url, []) :
{ id = p.id, precedence = p.precedence }],
# 10 of 87 apps in hb-infra production, not all of them
contains(app.access_policies, "vpn_bypass")
? [{ id = cloudflare_zero_trust_access_policy.thb_vpn_bypass.id, precedence = 3 }] : [],
)
}
}
resource "cloudflare_zero_trust_access_application" "apps" {
for_each = { for app in local.access_apps : app.domain_url => app }
policies = local.app_policies[each.key]
lifecycle {
precondition {
condition = (length(distinct([for p in local.app_policies[each.key] : p.precedence]))
== length(local.app_policies[each.key]))
error_message = "Policy precedences on ${each.key} must be unique."
}
}
}
Both built-in attachments stay conditional on the same
access_policies membership the v4 module uses today
(subdomains_with_email_policy,
subdomains_with_vpn_policy), so the migration does not
widen access. The precondition is what makes the #1431 class a plan-time
error: an external entry on 1 or 3, or two
external entries on the same value, fails before anything reaches the
API.
Written as a flat list keyed only by policy name —
var.external_policy_ids["botdojo"] with no application key
— BotDojo would be attached to every application in the
root. The same goes for an unconditional built-in: attaching
thb_vpn_bypass to every entry in
local.access_apps would add a VPN-IP bypass to roughly 77
applications that do not have one. That is the failure mode this example
exists to catch.
Place, do not manage. Any proposed input shape for external policies should be checked against it before being accepted: it has to express “precedence 2 on these two specific subdomains and nowhere else.” The sketch above is illustrative, not a proposed interface — the real shape depends on the Phase 1 spike findings.
4. Precedence becomes per-attachment data
Precedence does not vanish in v5 — it moves. Each entry in an
application’s policies list carries its own explicit
precedence value, so the module assembles a list of
{ id, precedence } references rather than relying on list
order. What goes away is precedence as a property of the
policy, and with it the need for any module-level
precedence variable: the values are computed per
application from the policies that application actually declares,
including the external ones from decision 3.
The interim v4 workaround does the crudest possible version of this:
terraform-modules#140
moves thb_vpn_bypass from precedence 2 to 3 and reserves 2
by convention, documented in the module README. It is a comment, not a
constraint — nothing stops the next hand-made policy landing on 1 or 3
and failing the same way. Replacing that convention with a declared list
is one of the concrete wins of migrating, and the v4 change is
deliberately cheap to discard because of it.
Risks
| Risk | Severity | Note |
|---|---|---|
| Policy-less application denies all traffic mid-migration | High | Attach-before-detach; prove in development first. Applying raw
tf-migrate output does exactly this — use the two-step
pattern under Retiring the
legacy policies |
| v3→v4 tunnel rename recreates tunnels and churns DNS | High | Phase 0. state rm + import against the
same tunnel ID; targeted plan on one v3 root first |
| Shared nonprod/prod account, staging change hits prod | High | Design decision 2; settle before coding |
| Hand-made policies (BotDojo) silently detached post-migration | Medium | Design decision 3; a new failure mode v5 introduces |
| Pipeline credential cannot manage service tokens | Medium | auth.forbidden on the service-token endpoints. External
policies can be referenced but their tokens cannot be owned without
widening cloudflare_access_app_api_token. See the BotDojo
worked example |
| Provider API rate limits | Medium | Already bitten once — c9231111 added rps = 2,
retries = 8 to hb-infra prod for Argo Tunnel rate-limit CI
failures. Root-wide plans across the whole Cloudflare surface will hit
this harder |
| Resources the provider does not move automatically | Medium | Fall back to removed + import;
tf-migrate --dry-run in the Phase 1 spike shows whether any
exist |
| Ten roots on six module tags (six tunnel tags) | Low | Opt-in per root, so partial migration is a valid resting state. Phase 0 collapses this to one tag per module |
Proposed Phasing
Deliberately not estimated — costing depends on what the Phase 0 targeted plan and the Phase 1 spike turn up.
- Phase 0 — v4 normalization. Every consuming root on
provider v4.52.5,
cloudflare_accesson the current v4 tag, andcloudflare_tunnelonv0.1.42or later, with no tunnel recreates. Start with a targeted plan on one v3 root to see exactly what it wants to replace. This is independently valuable — it retirescloudflare_argo_tunnel, puts one module tag everywhere, and gets the terraform-modules#140 fix into every root once it is tagged — and it is a hard prerequisite for v5 either way. Tracked as its own ticket. - Phase 1 — spike.
tf-migrate --dry-runagainst the module and one development root; confirm theincludeshape forip. Then prove the two-step policy pattern on a throwaway app, including that detached app-scoped policies are garbage-collected. Nothing past this starts until it lands. - Phase 2 — module rewrite. New major of
cloudflare_accessagainst provider>= 5.19, incorporating design decisions 1–4. Published as a new tag; no existing consumer is forced to it. - Phase 3 — development roots.
hb-infraandbees-infradevelopment. Whole Cloudflare surface per root, includingcloudflare_tunneland inline records. Lowest stakes, and these roots sit on their own Cloudflare account (945798da…), so mistakes there cannot reach prod. That is a property of these two roots, not of “development”: ha-infra development points at the sharede6c92b47…account. - Phase 4 — non-production. First contact with the shared account. Validate design decision 2 in anger.
- Phase 5 — production. One root at a time, with the
ha-infraandhb-infraadmin surfaces last.
Open Questions
What are the v5 import ID formats?Largely moot: the in-scope types move viamovedon v5.19+. Still needed for anythingtf-migrate --dry-runshows is not moved automatically.- What is the exact v5
includeshape for IP lists — one object per CIDR, or a list inside one object? (blocks the module rewrite;tf-migrateoutput should answer it) DoesSuperseded bycf-terraformingcover these resource types?tf-migrate. What remains is generating the existing per-app policy UUIDs for step 1 of the two-step pattern.- Is there appetite to do this at all in the near term, given that terraform-modules#140 holds indefinitely? A deprecated-but-working provider is a legitimate resting state. Note the workaround buys a precedence slot by convention only — it does not remove the failure class, so “not yet” should be a decision about timing, not a belief that the problem is solved. Current answer (review): yes to the v4.52.5 normalization near term; “not yet” on v5 itself until that is done. The v3→v4 bump is where the downtime risk actually lives (the tunnel rename), so it gets its own ticket and a targeted plan on one v3 root first.
- Do we want the
cloudflare_tunnelmigration in the same RFC, since it is forced by the same provider bump? Yes — they cannot be separated per root, and its v3→v4 step is Phase 0.
Out of Scope
- Whether prod Green/BEA CRM admin should have a VPN bypass. That is an access-policy decision, independent of the provider version — since settled as yes. It comes back once terraform-modules#140 is merged and tagged and infrahive#1431 is reverted; neither has happened yet.
- Bringing BotDojo’s policy under Terraform before the migration. Ruled out on the evidence in the worked example, not deferred for convenience.
common-infra/shared, already on v5.- Any change to Cloudflare Access posture, session durations, or group membership. This migration is intended to be behaviour-preserving.