This error check was meant to terminate sendLoop and cause it to return
with the error from WriteTo. (Except for the special case where the
error is a net.Error that is also Temporary(), in which case we merely
logged the error and continued running sendLoop.)
Errors from WriteTo (whether Temporary() or not) were rare. I managed to
get one line this after several days' uptime on a server with heavy use:
sendLoop: write udp [::]:5300->X.X.X.X:YYYYY: sendto: operation not permitted
The above dnstt-server error was accompanied by a Linux kernel log
message:
nf_conntrack: nf_conntrack: table full, dropping packet
What happened is the conntrack table filled and failed to track the
state of some UDP exchanges. A UDP 4-tuple lost the RELATED state and
and outbound packet was blocked by the local firewall ("operation not
permitted").
This may not be the only way a non-Temporary() WriteTo error could
happen. But in any case, when one did happen, it would cause sendLoop to
return and the server to stop processing traffic. (Before
37129955de, this was especially bad,
because the return of sendLoop would not terminate the program: it would
keep running and receiving queries, but never send any responses. Now,
at least, the program terminates, so the failure is immediately
detectable.)
Now we simply log errors from WriteTo, as if they were always temporary.
The only exception is net.ErrClosed, which causes sendLoop to terminate
as before.
Background on the net.Error Temporary() pattern:
* "Use net.Error to distinguish temporary Accept errors."
https://gitlab.torproject.org/tpo/anti-censorship/pluggable-transports/goptlib/-/commit/3030f080eecf72b0e896236fca5fabd245c00bdb
* "Don't report errors that are not caused by Accept in AcceptSocks."
https://gitlab.torproject.org/tpo/anti-censorship/pluggable-transports/goptlib/-/commit/50b39b746c6ff34bf31977b658848d876ee84fbf
* https://go.dev/blog/error-handling-and-go#the-error-type
net.Error.Temporary was deprecated in go1.18:
* "net: deprecate Temporary error status"
https://github.com/golang/go/issues/45729
* https://go.dev/doc/go1.18#netpkgnet
See also:
* "net/http: server.Serve() uses deprecated net.Error.Temporary()"
https://github.com/golang/go/issues/66208
* "proposal: net: add ErrRetryableAcceptError"
https://github.com/golang/go/issues/66252
For now, though, even though I'm removing the Temporary() check on
WriteTo errors, I'm keeping it for KCP AcceptKCP and AcceptStream. It
may still be the right thing for an accept loop; cf.
https://groups.google.com/g/golang-nuts/c/-JcZzOkyqYI/m/wp_5G8LmAwAJ:
While the whole suite of Temporary errors isn't really coherent,
the issue is that a small subset of Temporary is still useful
for Accept loops and it doesn't have a non-deprecated
replacement. As a case in point, I presume that http.Server is
going to keep using Temporary indefinitely.
Without sendLoop, the program can no longer make progress. Treat
sendLoop and recvLoop as peer goroutines, instead of having recvLoop on
the main class stack and sendLoop as a goroutine.
Before go1.23, calling Stop, and draining the channel if the timer did
not already fire, is necessary before calling Reset:
https://pkg.go.dev/time@go1.22.12#Timer.Reset
This changed in go1.23: now Reset automatically effectively drains the
channel, and calling Stop is no longer necessary.
https://pkg.go.dev/time@go1.23.9#Timer.Reset
However, the changes in go1.23 only take effect if go.mod specifies
1.23 or later. We currently specify 1.21.
https://go.dev/doc/go1.23#timer-changes
For compatibility, do the Stop/drain procedure before calling Reset. We
were already doing this for pollTimer in DNSPacketConn) sendLoop in
dnstt-client.
This change may not have any observable effect. The duration we Reset
the timer to was 0, so if there had been a stale value in the channel
because of a failure to drain it, the effect would be the same as
waiting 0 seconds. We were already calling Stop when finished with the
timer, so it would have been garbage-collectable even before go1.23.
Chrome fingerprints appear to be usable now, as long as they are not the
oldest ones.. (Compare to the commit log message of
98bdffa1706dfc041d1e99b86c47f29d72ad3a0c.) Also add the "random"
fingerprint.
-doh dns.google -dot dns.google -doh 1.1.1.1 -dot 1.1.1.1
none ok ok ok ok
random ok ok ok ok
Firefox_55 ok ok ok ok
Firefox_56 ok ok ok ok
Firefox_63 ok ok ok ok
Firefox_65 ok ok ok ok
Firefox_99 ok ok ok ok
Firefox_102 ok ok ok ok
Firefox_105 ok ok ok ok
Firefox_120 ok ok ok ok
Chrome_58 ERROR ERROR ok ok
Chrome_62 ERROR ERROR ok ok
Chrome_70 ERROR ERROR ok ok
Chrome_72 ok ok ok ok
Chrome_83 ok ok ok ok
Chrome_87 ok ok ok ok
Chrome_96 ok ok ok ok
Chrome_100 ok ok ok ok
Chrome_102 ok ok ok ok
Chrome_120 ok ok ok ok
iOS_11_1 ok ok ok ok
iOS_12_1 ok ok ok ok
iOS_13 ok ok ok ok
iOS_14 ok ok ok ok
This is the script used to collect data for the above table:
TCPDUMP="tcpdump"
FPS="
none random Firefox_55 Firefox_56 Firefox_63 Firefox_65
Firefox_99 Firefox_102 Firefox_105 Firefox_120 Chrome_58
Chrome_62 Chrome_70 Chrome_72 Chrome_83 Chrome_87 Chrome_96 Chrome_100
Chrome_102 Chrome_120 iOS_11_1 iOS_12_1 iOS_13 iOS_14
"
sudo -v; \
for HOST in dns.google 1.1.1.1; do
for UTLS in $FPS; do
ID="doh-utls-$HOST-$UTLS";
echo
echo "$ID"
sudo $TCPDUMP -n -U -w "$ID.pcap" &
sleep 1;
timeout 2 ./dnstt-client -doh https://"$HOST"/dns-query -utls "$UTLS" -pubkey-file "$PUBKEY" "$DOMAIN" 127.0.0.1:7000;
sudo kill $!;
tshark -n -V -Y ssl.handshake.ciphersuites -r "$ID.pcap" | sed -n -e '/^Transport Layer Security/,/^$/p' > "$ID.txt";
ID="dot-utls-$HOST-$UTLS";
echo
echo "$ID"
sudo $TCPDUMP -n -U -w "$ID.pcap" &
sleep 1;
timeout 2 ./dnstt-client -dot "$HOST":853 -utls "$UTLS" -pubkey-file "$PUBKEY" "$DOMAIN" 127.0.0.1:7000;
sudo kill $!;
tshark -n -V -Y ssl.handshake.ciphersuites -r "$ID.pcap" | sed -n -e '/^Transport Layer Security/,/^$/p' > "$ID.txt";
done;
done
Remove the SNI workaround that is fixed in v1.6.6.
https://github.com/refraction-networking/utls/issues/96
I find that the workaround for
https://github.com/refraction-networking/utls/issues/75
is still necessary; otherwise the local side thinks it's speaking
HTTP/1.1 while the remote is speaking HTTP/2:
2024/05/13 17:15:49 sendLoop: Post "https://1.1.1.1/dns-query": net/http: HTTP/1.x transport connection broken: malformed HTTP response "\x00\x00\x12\x04\x00\x00\x00\x00\x00\x00\x03\x00\x00\x00d\x00\x04\x00\x01\x00\x00\x00\x05\x00\xff\xff\xff\x00\x00\x04\b\x00\x00\x00\x00\x00\x7f\xff\x00\x00\x00\x00\b\a\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x01"
In testing with uTLS v1.0.0, go1.15.15, -doh and -dot, and dns.google
and 1.1.1.1, there is no Chrome fingerprint that works on all of them. I
did not investigate what exactly is going wrong. The error message
generally is "remote error: tls: unexpected message".
-doh dns.google -dot dns.google -doh 1.1.1.1 -dot 1.1.1.1
Firefox_55 ok ok ok ok
Firefox_56 ok ok ok ok
Firefox_63 ok ok ok ok
Firefox_65 ok ok ok ok
Chrome_58 ERROR ERROR ok ok
Chrome_62 ERROR ERROR ok ok
Chrome_70 ERROR ERROR ERROR ok
Chrome_72 ok ok ERROR ok
Chrome_83 ok ok ERROR ok
iOS_11_1 ok ok ok ok
iOS_12_1 ok ok ok ok
This is a script I used for testing fingerprints:
FPS="none Firefox_55 Firefox_56 Firefox_63 Firefox_65 Chrome_58 Chrome_62 Chrome_70 Chrome_72 Chrome_83 iOS_11_1 iOS_12_1"
sudo -v; \
for HOST in dns.google 1.1.1.1; do \
for UTLS in $FPS; do \
ID="doh-utls-$HOST-$UTLS"; \
sudo tcpdump -n -U -w "$ID.pcap" & \
timeout 2 ./dnstt-client -doh https://"$HOST"/dns-query -utls "$UTLS" -pubkey-file "$PUBKEY" "$DOMAIN" 127.0.0.1:7000; \
sudo kill $!; \
tshark -n -V -Y ssl.handshake.ciphersuites -r "$ID.pcap" | sed -n -e '/^Transport Layer Security/,/^$/p' > "$ID.txt"; \
ID="dot-utls-$HOST-$UTLS"; \
sudo tcpdump -n -U -w "$ID.pcap" & \
timeout 2 ./dnstt-client -dot "$HOST":853 -utls "$UTLS" -pubkey-file "$PUBKEY" "$DOMAIN" 127.0.0.1:7000; \
sudo kill $!; \
tshark -n -V -Y ssl.handshake.ciphersuites -r "$ID.pcap" | sed -n -e '/^Transport Layer Security/,/^$/p' > "$ID.txt"; \
done; \
done