Skip to content

feat(iroh): reach IPv4 addresses through NAT64 on IPv6-only networks - #4577

Draft
huangyingw wants to merge 1 commit into
n0-computer:mainfrom
huangyingw:nat64
Draft

huangyingw wants to merge 1 commit into
n0-computer:mainfrom
huangyingw:nat64

Conversation

@huangyingw

Copy link
Copy Markdown

Description

Closes #4576.

IPv6-only endpoints behind NAT64 (most large mobile carriers, e.g. T-Mobile US) can't get a direct path to IPv4-only remotes today: they share no address family, so the connection stays on the relay. This adds NAT64 support so they can.

The translation happens in the IP transports, like a userspace CLAT (RFC 6877):

  • While a NAT64 prefix is active, datagrams to public IPv4 destinations are sent from the IPv6 socket to the RFC 6052 synthesized address; datagrams received from inside the prefix are reported as coming from the embedded IPv4 address. noq, path management and the remote never see the IPv6 form.
  • IPv4 QAD uses the same sockets, so it then succeeds through NAT64 and returns the NAT64 gateway's public address. That is published as a reflexive candidate, which lets remotes behind EIM NATs hole punch too — not only remotes with a reachable port.

When to translate is decided in net_report (Client::update_nat64):

  • IPv4 QAD failed and the host has an IPv6 address, and either
    • the network's DNS64 reveals the prefix (ipv4only.arpa, RFC 7050), or
    • the host has no IPv4 address other than a CLAT one (192.0.0.0/29), in which case the Well-Known Prefix 64:ff9b::/96 is assumed. iOS can hold a CLAT address on some interface without offering IPv4 to apps, so plain have_v4 isn't enough here.
  • A failed IPv4 QAD on a dual-stack host without DNS64 never diverts native IPv4.
  • The host's IPv6 address is used rather than udp_v6, since a relay may not offer IPv6 QAD even where IPv6 works.
  • Once on, translation stays on until the next major network change: from then on IPv4 QAD succeeds through NAT64, so udp_v4 no longer measures native IPv4.

Changes:

  • net_report/nat64.rs (new): RFC 6052 synthesis/extraction for all six prefix lengths, RFC 7050 prefix discovery, is_translatable, and Nat64State shared between net_report and the transports (an AtomicBool fast path, so nothing changes on the hot path when NAT64 is off). Hand-written because the rfc6052 crate is GPL-3.0.
  • net_report.rs: update_nat64 / discover_nat64_prefix; IPv4 QAD also runs while translation is active.
  • net_report/reportgen.rs: IfStateDetails::have_native_v4 (IPv4 excluding the CLAT range).
  • socket/transports.rs, socket/transports/ip.rs: translate on send, map back on receive.
  • tests/patchbay/nat64.rs (new): IPv6-only client behind IspV6 ↔ IPv4-only server, with PublicV4 and with Home.

API Changes

  • Report::nat64_prefix: Option<Ipv6Net> (Report is #[non_exhaustive], so this is additive). It also reflects that udp_v4 / global_v4 describe IPv4 through NAT64 while it is set.
  • ipnet now enables its serde feature (for the new Report field).

Notes & open questions

  • Tests. The two new patchbay tests fail on main (no direct path within 30s, relay only) and pass with this change. Full patchbay suite on this branch vs main, test by test: the only difference is the two new tests (fail on main, pass here); all other tests have the same result (63 pass, 9 ignored). degrade_client_very_bad_cellular failed once in the full run on this branch while the machine was under heavy load, then passed 5/5 when rerun on its own (also 5/5 on main). iroh lib tests: 134 passed. cargo fmt and cargo clippy --all-targets are clean.
  • Real devices. Not yet verified on a real T-Mobile iPhone; I plan to do that next and will report back here.
  • DNS64 on mobile. Whether the resolver iroh uses on iOS/Android actually goes through the carrier's DNS64 is something I couldn't verify. The CLAT-aware Well-Known-Prefix fallback covers the T-Mobile case regardless, but an OS-provided prefix (e.g. iOS NAT64 prefix APIs, RFC 8781 PREF64) might be a better source. Happy to hear your preference.
  • Network-change handling relies on is_major. If there's a better signal for "the IPv4 situation may have changed", I'd switch to it.
  • Disclosure. This change was written with an AI coding assistant and validated with the tests above; I've left the "created by a human" checklist item unchecked for that reason.

Change checklist

  • Self-review.
  • Documentation updates following the style guide, if relevant.
  • Tests if relevant.
  • All API changes documented.
  • This PR was created by a human that thought critically about the
    proposed change and wrote an as clear and concise description as
    they could.
  • This PR isn't slop, and is carefully crafted to do have the
    intented effect.

…works

On IPv6-only networks with NAT64 (most large mobile carriers, e.g.
T-Mobile US), an endpoint has no usable IPv4 socket. When the remote only
has IPv4 addresses, the two sides share no address family and the
connection stays on the relay forever, even when the remote's port is
directly reachable.

This translates transparently in the IP transports, like a userspace CLAT
(RFC 6877): while a NAT64 prefix is active, datagrams to public IPv4
destinations are sent from the IPv6 socket to the RFC 6052 synthesized
address, and datagrams from inside the prefix are reported as coming from
the embedded IPv4 address. noq, path management and the remote never see
the IPv6 form.

Since the relay's QAD probes use the same sockets, IPv4 QAD then succeeds
through NAT64 and yields the NAT64 gateway's public address, which is
published as a reflexive candidate. This lets the remote hole punch too,
so remotes behind an EIM NAT (typical home routers) become reachable as
well, not only remotes with a reachable port.

net_report decides when to translate: IPv4 QAD failed while the host has
IPv6, and either the network's DNS64 reveals the prefix (RFC 7050,
ipv4only.arpa), or the host has no IPv4 address besides a CLAT one
(192.0.0.0/29), in which case the Well-Known Prefix is assumed. A failed
IPv4 QAD on a dual-stack host without DNS64 never diverts native IPv4.
Once on, translation stays on until the next major network change, since
IPv4 QAD then succeeds through NAT64 and no longer measures native IPv4.
The prefix is exposed as `Report::nat64_prefix`.

New patchbay tests put an IPv6-only client behind the `IspV6` (NAT64)
preset and connect to an IPv4-only server, once with a public address and
once behind a Home NAT. Both fail on main (relay only, no direct path
within 30s) and pass with this change.

Claude-Session: https://claude.ai/code/session_01YRX29ic8ZvvDHLKfF3En1q

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: 🚑 Needs Triage

Development

Successfully merging this pull request may close these issues.

IPv6-only endpoints behind NAT64 never get a direct path to IPv4-only endpoints

1 participant