@hugo@social.treehouse.systems
2026-08-27 05:21 UTC
K, need the #ipv6 gang.
I'm hanging on the edge of being That Guy here.
The k8s network model always seemed to me to map really nicely onto a simple /64 per worker. Containers in a pod share a network namespace, not allocating discrete IP addresses to individual containers.
But, reading the proposed draft for DC v6 deployment, and specifically the section at https://www.ietf.org/archive/id/draft-martin-deploying-ipv6-data-center-02.html#name-prefix-allocation-for-hosts re container hosts, and taking some pause here:
A common data center pattern assigns a /56 to each physical host (or rack entity), providing 256 /64 subnets --- one /64 for the host itself and up to 255 /64 prefixes for containers, virtual machines, or Kubernetes pods.
Does it? Like, is that actually "common"? I don't want to just port IPv4 thinking over here, but I honestly struggle to think of scenarios where I would specifically want a /64 per pod.
Assigning only a /64 per host (or per rack entity without further delegation) is often insufficient when multiple containers each need their own address space.
Eh? That...seems to run counter to the common k8s network model? How often are folks finding this "insufficient"? Are we running a bunch of weird VNF workloads I'm not yet encountering?
In a closed data center with explicit routing and no SLAAC on container segments, some designs assign one /64 per physical host and carve /72 (or longer) subnets from that host prefix for container tiers. That pattern is not suitable on the public Internet or where hosts expect standard /64 semantics; use it only with operator-wide agreement and tested CNI or orchestrator support.
Maybe I'm missing something, but any CNI patterns I've come across so far seem to basically expect an address per pod, not a prefix per pod with multiple discrete container addresses within that. Who's running SLAAC between the worker and containers? Are we the exception here with "explicit routing"?
Honestly I was originally perhaps a bit conservative with just an address for the worker itself in a rack-local uplink /64, and then a /64 per worker for pods. Perhaps a /64 or two for the worker itself and then a /64 for its pods is more suitable. But a /56 per worker to support a /64 per pod seems...a bit much?
A /52 podCIDR supernet gives us 4,096x /64s, so 4k workers at a /64 per worker. That's still a good chunk to work with if we have big clusters. Shifting to a /56 per worker would put us at a /44 podCIDR supernet for a 4k node cluster limit.
I know we should work from our base address needs and then work out way up rather than trying to "fit" within a predefined sizing. But this seems like a good chunk of hierarchy bits that don't actually fit with my understanding of the intended (k8s) network model.
Are folks really out here dropping a /64 per k8s pod?
Replies (1)
-
@karppinen@mastodon.online 2026-08-27 05:39
@hugo@social.treehouse.systems I’ll just take this opportunity to rant about ARIN policies. Over here in RIPE land where everyone gets a /29 all that sounds very reasonable. Then ARIN comes in and decides that folks who want more than a /48 should be paying extra. Artificial scarcity that guarantees folks will remain v4-brained.