Commit Graph
64 Commits
Author SHA1 Message Date
David Fifield fec00a2a78 -utls option and random TLS fingerprint selection. 2022-01-02 19:16:00 -07:00
David Fifield 70c78e24d7 TLS camouflage using uTLS and a hardcoded Client Hello ID.
net/http Transport.DialTLSContext requires go1.14.
https://go.dev/doc/go1.14#net/http
2022-01-02 19:14:35 -07:00
David Fifield d92a791b68 Don't leave TLSPacketConn unclosed when there's a first Dial error. 2021-12-24 07:37:24 -07:00
David Fifield e4dc2883ef Use errors.Is to compare against ErrClosedPipe.
I was still getting "io: read/write on closed pipe" errors in the logs,
even after comparing errors against io.ErrClosedPipe to skip logging
them. It turns out that kcp-go wraps many of its errors in another type.
The actual type of the errors was *errors.withStack, where errors is
https://github.com/pkg/errors. We can use the go1.13 errors interface
(https://blog.golang.org/go1.13-errors) to get at the value inside.
2021-08-03 21:00:45 -06:00
David Fifield 1f73f6f5b6 Ignore ErrClosedPipe in "copy stream←upstream" as well.
Saw this happen on the server during the 2021-08-02 performance tests.
Doing on the client, too, for uniformity.
2021-08-03 20:58:20 -06:00
David Fifield de15c5a512 Performance tuning: MaxStreamBuffer, SetWindowSize, QueueSize.
This enlarges a few buffers and windows, with the goal of improving
download performance. kcp's SetWindowSize controls the number of
unacknowledged packets that are allowed. smux's MaxStreamBuffer is
another kind of "receive window" that is advertised to the peer of how
much we are willing to receive at once. The default MaxStreamBuffer is
64 KB, but kcptun overrides the default to 2 MB. turbotunnel's QueueSize
is the size of internal buffers in QueuePacketConn and RemoteMap;
empirically I found that the server would sometimes fill its outgoing
buffer if SetWindowSize and QueueSize were equal, so I set QueueSize to
be twice SetWindowSize.

https://lists.torproject.org/pipermail/anti-censorship-team/2021-July/000178.html
https://gitlab.torproject.org/tpo/anti-censorship/pluggable-transports/snowflake/-/merge_requests/48

The changes have a large effect on a direct -udp connection without a
recursive resolver—which, however, is a discouraged configuration.
Through a recursive resolver, the improvements are more modest. If I
really crank up the buffer sizes, I can get surprisingly fast downloads
over a direct -udp connection (over 1 MB/s), but a connection through a
resolver doesn't keep getting faster and may even get slower. I want to
avoid a bufferbloat situation with oversized buffers, too. I manually
explored a small neighborhood of parameter values and picked some
settings that looked reasonable.

The tables below show the test results. The test is downloading 10 MiB
between two servers with 100 ms RTT between them. Server:
	dnstt-server -udp :53 -privkey-file server.key t.example.com 127.0.0.1:9321
	ncat -l -k -v 9321 --send-only --sh-exec 'dd bs=1M count=10 if=/dev/urandom'
Client:
	dnstt-client -pubkey-file server.pub t.example.com 127.0.0.1:7000
	ncat --recv-only 127.0.0.1 7000 | pv -t -r -a -b -i 0.2 > /dev/null
I did the download under every treatment twice and recorded the download
rate in KiB/s. "Server drops" comes from hacking some log messages to
turbotunnel.QueuePacketConn to track how often the "Drop the incoming
packet" (QueueIncoming method) and "Drop the outgoing packet" (WriteTo)
cases happen.

resolver	method	QueueSize	MaxStreamBuffer	SetWindowSize	KiB/s	KiB/s
--------	------	---------	---------------	-------------	-----	-----
direct		udp	64		64*1024		(32, 32)	 169	 173	(status before this commit)
dns.google	udp	64		64*1024		(32, 32)	  63.8	  64.3	(status before this commit)
dns.google	doh	64		64*1024		(32, 32)	 125	 122	(status before this commit)

resolver	method	QueueSize	MaxStreamBuffer	SetWindowSize	KiB/s	KiB/s
--------	------	---------	---------------	-------------	-----	-----
direct		udp	64		1*1024*1024	(32, 32)	 172	 174
dns.google	udp	64		1*1024*1024	(32, 32)	  57.3	  58.4	server drops
dns.google	doh	64		1*1024*1024	(32, 32)	 128	 128

resolver	method	QueueSize	MaxStreamBuffer	SetWindowSize	KiB/s	KiB/s
--------	------	---------	---------------	-------------	-----	-----
direct		udp	64		1*1024*1024	(64, 64)	 322	 305
dns.google	udp	64		1*1024*1024	(64, 64)	  72.5	  70.9	server drops
dns.google	doh	64		1*1024*1024	(64, 64)	 136	 139	server drops

resolver	method	QueueSize	MaxStreamBuffer	SetWindowSize	KiB/s	KiB/s
--------	------	---------	---------------	-------------	-----	-----
direct		udp	128		1*1024*1024	(64, 64)	 321	 325	(this commit)
dns.google	udp	128		1*1024*1024	(64, 64)	  82.5	  78.5	(this commit)
dns.google	doh	128		1*1024*1024	(64, 64)	 129	 131	(this commit)

resolver	method	QueueSize	MaxStreamBuffer	SetWindowSize	KiB/s	KiB/s
--------	------	---------	---------------	-------------	-----	-----
direct		udp	2048		4*1024*1024	(1024, 1024)	1240	1060	server drops
dns.google	udp	2048		4*1024*1024	(1024, 1024)	  73.5	 81.4
dns.google	doh	2048		4*1024*1024	(1024, 1024)	 115	 129
2021-08-02 15:25:19 -06:00
David Fifield 1251acc351 Simplify DNSPacketConn.sendLoop a little. 2021-08-02 01:00:16 -06:00
David Fifield c7613b89e1 Reduce smux idle timeout from 10 minutes to 2 minutes. 2021-08-01 22:32:57 -06:00
David Fifield 27dbee1b66 Make pollChan buffered, and send on it only once.
Formerly we sent twice on pollChan, but because it was unbuffered, the
second send was almost always dropped. KCP's own ACK packets should also
serve as another source of what are effectively polling queries at a
rate proportional to the rate at which we are receiving.

I did some performance tests of downloading 10 MiB between two servers
with 100 ms RTT between them. Server:
	dnstt-server -udp :53 -privkey-file server.key t.example.com 127.0.0.1:9321
	ncat -l -k -v 9321 --send-only --sh-exec 'dd bs=1M count=10 if=/dev/urandom'
Client:
	dnstt-client -pubkey-file server.pub t.example.com 127.0.0.1:7000
	ncat --recv-only 127.0.0.1 7000 | pv -t -r -a -b -i 0.2 > /dev/null
I did the download under every treatment twice and recorded the download
rate in KiB/s.

First, the results for the commit before this one. I also hacked in
fewer sends on pollChan for each packet received. The result for 1 poll
are about the same as for 2 polls, which is expected, with the
observation that the unbuffered pollChan was usually dropping the second
send. 0 polls results in a very slow rate (possibly driven only by smux
keepalive packets). ("~" means I stopped waiting for the download after
about 2 minutes.)

resolver	method	pollChan cap	SetACKNoDelay	polls	KiB/s	KiB/s
--------	------	------------	-------------	-----	-----	-----
direct		udp	unbuffered	false		2	146	159	(status quo before this commit)
dns.google	udp	unbuffered	false		2	 58.3	 60.2	(status quo before this commit)
dns.google	doh	unbuffered	false		2	125	126	(status quo before this commit)
direct		udp	unbuffered	false		1	155	165
dns.google	udp	unbuffered	false		1	 57.5	 58.7
dns.google	doh	unbuffered	false		1	124	123
direct		udp	unbuffered	false		0	 ~7	~11
dns.google	udp	unbuffered	false		0	 ~6	 ~5
dns.google	doh	unbuffered	false		0	 ~6	 ~6

Now, the result after this commit. A buffered pollChan with 1 poll per
receive has almost identical performance to an unbuffered pollChan with
1 or 2 polls per receive, which is expected. Using 2 polls rather than 1
actually helps performance a fair bit in the direct/UDP and Google/UDP
treatments, but hurts performance in the Google/DoH treatment. 0 polls
still yields poor performance.

resolver	method	pollChan cap	SetACKNoDelay	polls	KiB/s	KiB/s
--------	------	------------	-------------	-----	-----	-----
direct		udp	16		false		2	174	175
dns.google	udp	16		false		2	 76.0	 75.3
dns.google	doh	16		false		2	 68.7	68.0
direct		udp	16		false		1	149	151	(this commit)
dns.google	udp	16		false		1	 60.2	 60.3	(this commit)
dns.google	doh	16		false		1	123	121	(this commit)
direct		udp	16		false		0	 ~7	 ~8
dns.google	udp	16		false		0	 ~5	 ~5
dns.google	doh	16		false		0	 ~6	 ~7

I tried the additional modification of calling conn.SetACKNoDelay(true).
My guess was that this would cause every received data packet to be
ACKed immediately, which should have the same function as a poll.
Unexpectedly for me, SetACKNoDelay(true) actually slows down the
direct/UDP and Google/DoH cases a lot. But strangely, Google/UDP becomes
faster (and 1 poll is even faster than 2 polls in that case). Strangest
of all, SetACKNoDelay(true) makes the 0-poll Google/UDP and Google/DoH
treatments run reasonably fast. My best guess as to why that is the case
is that routing through a Google resolver tends to disorder the packet
sequence, which maybe results in more ACKs, which effectively act as
polls.

resolver	method	pollChan cap	SetACKNoDelay	polls	KiB/s	KiB/s
--------	------	------------	-------------	-----	-----	-----
direct		udp	16		true		2	 87.7	 84.2
dns.google	udp	16		true		2	 82.7	 68.9
dns.google	doh	16		true		2	 59.8	 60.8
direct		udp	16		true		1	 54.1	 47.8
dns.google	udp	16		true		1	 97.5	100
dns.google	doh	16		true		1	 71.4	 70.9
direct		udp	16		true		0	 ~9	 ~8
dns.google	udp	16		true		0	 48.0	 48.6
dns.google	doh	16		true		0	 59.2	 59.2
2021-08-01 22:09:03 -06:00
David Fifield 4de69201d1 Add a dial timeout to TLSPacketConn.
In my testing locally, specifying -dot with a non-responsive TCP port
would time out after about 30 seconds anyway:
	$ time ./dnstt-client -dot tns.example.com:8000 -pubkey-file server.pub t.example.com 127.0.0.1:7000
	dial tcp 45.79.134.119:8000: connect: connection timed out
	real	0m31.398s
	user	0m0.006s
	sys	0m0.003s
Which is in line with the documentation for net.Dialer:
	https://golang.org/pkg/net/#Dialer
	With or without a timeout, the operating system may impose its
	own earlier timeout. For instance, TCP timeouts are often around
	3 minutes.
But may as well be explicit.

This commit has the side effect of changing the error message from
"connection timed out" to "i/o timeout".
	$ time ./dnstt-client -dot tns.example.com:8000 -pubkey-file server.pub t.example.com 127.0.0.1:7000
	dial tcp 45.79.134.119:8000: i/o timeout
	real	0m30.007s
	user	0m0.003s
	sys	0m0.007s
I tried setting the dialTimeout to 40 seconds, and in that case the
system timeout take precedence after ≈31 seconds, with the "connection
timed out" error as before.

This is issue UCB-02-007 from the 2021 security audit of Turbo Tunnel by
Cure53.

The audit report additionally recommends calling SetReadDeadline before
each read operation. I have chosen not to do that. It is intended that
the TLS connection should be able to remain idle if there is nothing to
send. As DNS is a query–response protocol, one might expect a response
(and within a certain amount of time) only after sending a query;
sendLoop could refresh the ReadDeadline for recvLoop every time it sends
a query. But a malicious DoT server could keep a useless connection
alive anyway by sending Slowloris-style short responses within each
deadline, and an external adversary could capable of delaying responses
could deny service indefinitely or simply block the server. In any case,
the smux KeepAliveTimeout serves as a check that prevents stalled
connections from remaining indefinitely.
2021-04-20 17:32:15 -06:00
David Fifield 064c53e3d1 Be uniform about not ending log calls with "\n". 2021-04-20 15:14:10 -06:00
David Fifield bd8a0e870f Grammar uniformity "ourself"→"ourselves". 2021-04-20 13:55:38 -06:00
David Fifield 7b033a38ca Reflow comment. 2021-04-20 12:57:47 -06:00
David Fifield 58c01e740f requestor → requester
RFC 1035 uses "requester" and RFC 6891 uses "requestor". I think I
prefer "requester", and aspell agrees.
2020-08-30 19:50:23 -06:00
David Fifield e1e27bd4ce Comment typo. 2020-07-26 09:59:27 -06:00
David Fifield 3254c1c81e Open the client's local listener first. 2020-04-29 23:56:36 -06:00
David Fifield e48d53ceb0 Fix a log message. 2020-04-29 23:44:25 -06:00
David Fifield f6ca82d08c Buffer reads and writes in TLSPacketConn. 2020-04-29 12:44:07 -06:00
David Fifield ed2678917e Add a missing error check in TLSPacketConn.sendLoop. 2020-04-29 12:42:39 -06:00
David Fifield 4663433c08 Do receive-triggered polls based packets received.
Not amount of raw payload. This allows for the case where the received
payload is only padding, for example. (That can't happen with the
current downstream encoding scheme, which doesn't allow for padding, so
I believe this change results in equivalent behavior.)
2020-04-29 09:45:29 -06:00
David Fifield 05444dcb22 smux Stream.Write may also return EOF. 2020-04-25 21:27:14 -06:00
David Fifield 241225df1d Add -mtu option to server. 2020-04-25 20:54:09 -06:00
David Fifield 9ee6bf8abf Use wg.Add(2) instead of 2 × wg.Add(1). 2020-04-23 15:51:53 -06:00
David Fifield d14deab12b Documentation and light refactoring. 2020-04-19 17:16:27 -06:00
David Fifield 813a8564e8 Move some helper functions into dns.go. 2020-04-19 11:29:50 -06:00
David Fifield e9a98c3aef Avoid logging EOF and ErrClosedPipe errors.
smux Stream.WriteTo may return io.EOF, which breaks the contract of
io.Copy that says it should not return io.EOF. smux.Stream doesn't have
a unidirectional shutdown, so we always end up slamming it shut in both
directions and leave the other direction with a broken pipe.
2020-04-19 02:13:48 -06:00
David Fifield e7098959e2 Move global base32Encoding into dns.go. 2020-04-18 23:39:35 -06:00
David Fifield b7c18be90f Lowercase base32-encoded data in DNS names. 2020-04-18 23:39:35 -06:00
David Fifield 41c146355a Do MTU check first. 2020-04-18 19:05:23 -06:00
David Fifield 0a630f91c5 Log stream begin/end in server, log conv in client. 2020-04-18 19:05:23 -06:00
David Fifield 7ed79218eb Insert more padding when polling. 2020-04-18 18:46:54 -06:00
David Fifield b1bf9164a7 -privkey-file and -pubkey-file options. 2020-04-18 17:56:45 -06:00
David Fifield 505db3ec23 Server -privkey option and client -pubkey option. 2020-04-18 17:35:32 -06:00
David Fifield b98eb2c75e Move dummyAddr to turbotunnel.DummyAddr. 2020-04-18 16:02:14 -06:00
David Fifield 82ee14fefa Emit a log line even if we take no action on an unknown status code. 2020-04-18 16:02:14 -06:00
David Fifield b0f99e72bb Handle other unknown response status codes the same as 429. 2020-04-18 16:02:14 -06:00
David Fifield e501935578 Handle 429 Too Many Requests. 2020-04-18 16:02:14 -06:00
David Fifield 973f6310c5 Don't expire ClientMap if timeout is zero. 2020-04-18 16:02:14 -06:00
David Fifield 317c2c43d0 Don't send User-Agent.
HTTP header now looks like
```
POST / HTTP/1.1
Host: 127.0.0.1:8000
Content-Length: 221
Accept: application/dns-message
Content-Type: application/dns-message
Accept-Encoding: gzip
```
2020-04-18 16:02:14 -06:00
David Fifield 087d3b25dd Refactor HTTP response handling. 2020-04-18 16:02:14 -06:00
David Fifield e772e7bf2f Add a timeout to HTTP requests and follow redirects. 2020-04-18 16:02:14 -06:00
David Fifield 5350ea1e68 Increase initPollDelay to 500 ms. 2020-04-18 16:02:14 -06:00
David Fifield 7cbddf5fd9 Rework polling.
Give priority to data-carrying packets over polling packets.
Discard a polling packet whenever sending a data-carrying packet.
2020-04-18 16:02:14 -06:00
David Fifield 06144f76ae -dot mode. 2020-04-18 16:02:11 -06:00
David Fifield 1907daeba0 Simplify HTTPPacketConn. 2020-04-18 14:34:10 -06:00
David Fifield be9c3f1ac7 Refactor PacketConn handling. 2020-04-18 14:34:10 -06:00
David Fifield 7f3e9e4571 Overlay a noise layer atop KCP. 2020-04-18 14:34:10 -06:00
David Fifield 3f98c62207 Log when there's an error opening a stream. 2020-04-18 14:34:10 -06:00
David Fifield a6c891c5ae -doh mode. 2020-04-18 14:33:39 -06:00
David Fifield aefe4f9971 Factor out a pattern for different kinds of remote address. 2020-04-18 14:33:39 -06:00