TCP fundamentals and ChrisOS transfer path¶
About this chapter Frames, datagrams and streams
In this chapter
- TCP fundamentals and ChrisOS transfer path
- Scope
- Receive dispatch
- Minimum TCP header
- Generated data offset
- Flags
- Fixed advertised window
- Urgent data
- TCP checksum generation
- Receive checksum gap
- Common transmit builder
- Software-buffer size
- Null-payload contract
- Initial sequence numbers
- Built-in port-7 echo state
- Echo handshake
- Echo data
- Echo payload-bounds vulnerability
- Generic socket states
- Passive open
- Active open
- SYN retransmission
- Data transmit through sockets
- Remembered segment state
- Receive sequence behavior
- Socket receive bounds
- Receive-buffer pressure bug
- Blocking integration
- Connection close
- Peer FIN and RST
- Receive window and flow control
- Congestion control
- TCP options
- File-transfer port 9016
- CFS1 application framing
- Transfer validation
- Transfer TCP sequencing gap
- Transfer payload-bounds vulnerability
- Socket ownership evidence
- End-to-end evidence
- Concurrency
- Recommended hardening
- Current limitations
- Revision note
Scope¶
ChrisOS implements a small subset of TCP sufficient for a few project workflows, but it does not contain one complete RFC-style TCP engine.
There are three distinct TCP state machines:
- a single built-in echo connection on port 7 in
net.c; - the generic socket table in
sock.c; - a dedicated file-transfer connection on port 9016 in
net_xfer.c.
All three share the same packet builder and transmit checksum implementation, but their connection state, receive logic, timeout behavior, and application semantics are separate.
This chapter documents the implemented behavior rather than treating those paths as a full TCP implementation.
Receive dispatch¶
IPv4 protocol 6 enters net_tcp.
After minimum IPv4/TCP checks, dispatch order is:
destination port 9016
-> net_xfer_tcp()
otherwise
-> sock_on_tcp()
if consumed, stop
otherwise destination port 7
-> built-in echo TCP
otherwise
-> ignore
The transfer service therefore has priority over the generic socket table on its dedicated port.
Minimum TCP header¶
ChrisOS-generated TCP segments always use a 20-byte header with no TCP options.
The layout is conventional:
| Offset | Size | Field |
|---|---|---|
| 0 | 2 | source port |
| 2 | 2 | destination port |
| 4 | 4 | sequence number |
| 8 | 4 | acknowledgment number |
| 12 | 1 | data offset + reserved |
| 13 | 1 | flags |
| 14 | 2 | window |
| 16 | 2 | checksum |
| 18 | 2 | urgent pointer |
net_tcp_xmit writes all fields directly into the shared packet buffer.
Generated data offset¶
Transmit byte 12 is always:
The high nibble 5 means a five-word TCP header, or 20 bytes.
ChrisOS therefore generates no TCP options such as MSS, window scale, timestamps, or SACK-permitted.
Receive paths read the incoming data-offset nibble and allow headers larger than 20 bytes, but option contents are skipped rather than interpreted.
Flags¶
The code defines the classic low flag bits:
- FIN = 0x01;
- SYN = 0x02;
- RST = 0x04;
- PSH = 0x08;
- ACK = 0x10.
Generated traffic uses combinations such as SYN, SYN|ACK, ACK, PSH|ACK, and FIN|ACK.
There is no ECE/CWR handling and no NS support.
Fixed advertised window¶
Every segment generated by net_tcp_xmit advertises:
The peer's received window field is not used to throttle transmission.
There is no window scaling, zero-window handling, persist timer, or dynamic receive-window advertisement based on actual socket buffer space.
This is particularly important because each socket has only 2048 bytes of receive storage.
Urgent data¶
The urgent pointer generated by ChrisOS is always zero.
URG is not defined or handled by the current state machines.
Urgent data semantics are therefore unsupported.
TCP checksum generation¶
Transmit TCP checksum is calculated over the standard IPv4 pseudo-header plus TCP header and payload.
The pseudo-header contains:
- source IPv4 address;
- destination IPv4 address;
- zero byte;
- protocol 6;
- 16-bit TCP length.
The implementation then sums big-endian 16-bit words, includes a high-byte-padded final odd byte when necessary, folds carry, and returns one's complement.
net_tcp_xmit writes zero to the checksum field first, copies payload, computes the checksum, and writes the result at TCP offset 16.
Receive checksum gap¶
No receive TCP path validates the TCP checksum.
The built-in echo, socket table, and transfer service all trust segments without checking the pseudo-header checksum.
A corrupted or forged segment can therefore enter connection state processing if structural checks pass.
This is one of the highest-priority validation gaps.
Common transmit builder¶
net_tcp_xmit builds:
The IPv4 source is always the fixed ChrisOS address.
TCP sequence, acknowledgment, ports, and flags are supplied by the caller.
The function has no return value, so callers cannot determine whether the final link transmission succeeded.
Software-buffer size¶
The builder accepts packets while:
With 20-byte IPv4 and TCP headers, this permits up to 1546 bytes of TCP payload in the software packet buffer.
However, virtio_net_tx rejects Ethernet frames above 1514 bytes, giving an effective TCP payload ceiling of 1460 bytes for a normal header.
The upper layer can therefore construct a 1461..1546-byte payload that the driver silently rejects through the current void send wrapper.
Null-payload contract¶
net_tcp_xmit copies payload bytes whenever payload_len > 0.
It does not explicitly reject a null payload pointer with nonzero length.
Current internal callers generally pass valid storage, but the function contract is not hardened against invalid kernel callers.
Initial sequence numbers¶
The built-in echo server, socket connections, and transfer service derive initial sequence values from:
This is deterministic relative to the tick counter and is not a cryptographically unpredictable TCP ISN generator.
Connections created in related timing conditions may have correlated initial sequences.
The design is adequate for the project environment but not hardened against off-path sequence attacks.
Built-in port-7 echo state¶
net.c has one global TcpConn g_tcp.
It stores:
- active flag;
- remote MAC;
- remote IPv4 address;
- remote port;
- send-next;
- receive-next.
Only one echo connection can be represented.
A new accepted SYN overwrites the previous global state.
Echo handshake¶
For port 7, a segment containing SYN without ACK starts a connection.
ChrisOS:
- marks
g_tcp.active; - records remote MAC/IP/port;
- creates its initial sequence number;
- sets
rcv_nxt = incoming_seq + 1; - sends SYN|ACK;
- increments
snd_nxtby one.
A later ACK-only segment causes only the log message tcp established.
The received ACK number is read but not validated against snd_nxt.
Echo data¶
For the active remote IP/port, a PSH segment with payload is echoed back.
The implementation sets:
then sends ACK|PSH containing the same bytes and increments snd_nxt by payload length.
It does not require incoming_seq == previous rcv_nxt.
Out-of-order, duplicate, or overlapping data is therefore not handled according to full TCP stream semantics.
Echo payload-bounds vulnerability¶
The echo path derives payload length from the IPv4 total-length field.
It does not require that:
before passing frame + payload_off and payload_len to the transmit copy.
A forged IP total length larger than the physically received frame can therefore cause an out-of-bounds read while echoing data.
The IPv4 chapter identifies the same cross-layer length problem.
Generic socket states¶
The socket layer implements only four states:
- FREE;
- LISTEN;
- SYN_SENT;
- ESTABLISHED.
There is no explicit SYN_RECEIVED, FIN_WAIT_1, FIN_WAIT_2, CLOSE_WAIT, LAST_ACK, CLOSING, or TIME_WAIT.
This simplified state model affects both handshake timing and connection teardown.
Passive open¶
A LISTEN socket is matched by local destination port.
When a SYN without ACK arrives and no existing connection matches, a child socket is allocated.
The child is immediately marked ESTABLISHED, associated with the listener through parent, and sent SYN|ACK.
The implementation does not wait for the peer's final ACK before considering the child established.
sock_accept can therefore return that child before a full three-way handshake has been confirmed.
Active open¶
sock_connect allocates a socket, sets SYN_SENT, chooses an ephemeral local port starting at 40000, sets an initial sequence, and tries to send SYN.
If gateway MAC is unavailable, the helper cannot transmit even though the descriptor remains allocated.
On a received segment where:
the socket sets rcv_nxt = seq + 1, changes to ESTABLISHED, learns the source MAC, and sends ACK.
The received acknowledgment number is not checked against the locally expected sequence.
SYN retransmission¶
sock_tick is called from net_poll.
A socket remaining in SYN_SENT for more than 30 ticks retransmits its SYN using the original sequence.
There is:
- no maximum retry count;
- no exponential backoff;
- no connect timeout state;
- no error delivery to the caller.
The retry can continue indefinitely while polling continues.
Data transmit through sockets¶
sock_send requires ESTABLISHED state.
One call is limited to 200 payload bytes.
The segment is emitted with PSH|ACK.
The helper advances snd_nxt immediately by the number of payload bytes requested.
There is no send queue waiting for acknowledgment.
There is also no retransmission timer for application data.
Remembered segment state¶
The socket structure contains a 256-byte last array, last_len, and last_tick.
remember stores flags plus at most 250 bytes of the last payload.
However, the current sock_tick retransmission logic only resends SYN while state is SYN_SENT.
The remembered ordinary data is not used for a general TCP retransmission algorithm in the inspected revision.
Receive sequence behavior¶
When socket TCP payload passes the physical-bounds check, the code sets:
It does not compare the incoming sequence number to the previous expected rcv_nxt.
There is no out-of-order segment queue, duplicate detection, overlap trimming, or cumulative receive reassembly.
The receiver effectively trusts each accepted payload's sequence position.
Socket receive bounds¶
The generic socket path computes payload length from IPv4 total length but only copies when:
This is safer than the built-in echo and transfer paths.
It still depends on upstream validation for several malformed-header cases, and it does not verify TCP checksum.
Receive-buffer pressure bug¶
Each socket has a 2048-byte byte array.
If an incoming payload is larger than remaining room, only the portion that fits is copied.
Despite truncation, rcv_nxt advances by the complete incoming payload length and an ACK is transmitted.
The peer is therefore told that bytes were received even when the tail was discarded.
This violates reliable-stream semantics under receive-buffer pressure.
Blocking integration¶
For a native process owner, successful payload delivery calls proc_unblock(owner).
For CLVM receive syscall 144, an empty receive causes the VM to enter CLVM_WAITING, blocks the current process with the socket-block reason, and arranges to retry the syscall later.
The CLVM temporary buffer limits one send or receive syscall to 200 bytes.
This matches the socket send cap.
Connection close¶
sock_close sends FIN|ACK when the socket is ESTABLISHED.
Immediately afterward, it marks the socket FREE.
There is no FIN_WAIT state and no waiting for peer ACK or peer FIN.
Incoming FIN is not used to transition an established socket to a close state.
RST is also not processed as a connection abort.
The teardown implementation is therefore a best-effort notification rather than TCP close-state handling.
Peer FIN and RST¶
The flag constants exist, but ordinary socket receive logic has no dedicated FIN or RST branches.
The built-in echo and transfer service likewise lack complete teardown handling.
A peer closing a connection does not drive the classic TCP state machine.
This can leave application-visible semantics inconsistent with the remote endpoint.
Receive window and flow control¶
The peer's advertised TCP window is ignored.
ChrisOS also always advertises 65535 even if only a few bytes remain in the socket RX buffer.
There is no backpressure based on local capacity.
Combined with the truncation-and-ACK behavior, this means flow control is not currently reliable.
Congestion control¶
There is no congestion window, slow start, congestion avoidance, RTT estimator, retransmission timeout calculation, fast retransmit, or fast recovery.
Transmission is driven directly by application calls and the single SYN retry timer.
The current code should not be described as a congestion-controlled TCP stack.
TCP options¶
Incoming headers can have a data offset greater than 5, but the bytes are only skipped.
MSS is not negotiated.
Window scaling, timestamps, SACK, and other options are not implemented.
Generated SYN/SYN+ACK segments therefore contain no MSS option, leaving the peer to apply default behavior.
File-transfer port 9016¶
TCP destination port 9016 bypasses the socket table and enters net_xfer_tcp.
Only one global transfer connection exists.
A new SYN resets any previous transfer state, records the remote endpoint, sends SYN|ACK, and moves the application protocol to its header phase.
Like the echo server, it does not maintain a complete TCP handshake state.
CFS1 application framing¶
The host utility tools/cfs_send.py connects over TCP and sends:
The fixed application header is 12 bytes.
The guest consumes this stream through phases HDR, PATH, DATA, and DONE.
This application-level parser can consume one logical transfer across multiple sequential TCP payloads.
Transfer validation¶
The transfer service checks:
- four-byte CFS1 magic;
- path length nonzero and below 512;
- file size nonzero;
- file size at or below the filesystem maximum;
- allocation success.
The complete file is buffered with kmalloc, then written through fs_write.
Success sends byte 0x06; failure sends 0x15.
The host utility waits up to 30 seconds at the Python socket level for the one-byte application acknowledgment.
Transfer TCP sequencing gap¶
net_xfer_tcp sets rcv_nxt = seq + payload_length for each payload and consumes bytes immediately.
It does not require in-order sequence numbers or suppress retransmitted payload.
A duplicated TCP segment could therefore feed duplicate bytes into the CFS1 application state.
Out-of-order delivery can corrupt phase parsing.
The application parser handles segmentation, but the TCP layer does not provide reliable stream reassembly.
Transfer payload-bounds vulnerability¶
Like the built-in echo path, the transfer handler derives payload length from IPv4 total length without first requiring that the endpoint fit inside the physical received frame.
It then passes the resulting range to xfer_consume.
A forged oversized IPv4 total length can therefore produce an out-of-bounds read.
This path needs the same centralized length validation as the rest of TCP.
Socket ownership evidence¶
tools/test_sock_owner.c is the principal host-side socket test found in this revision.
It verifies:
- one CLVM slot cannot close another slot's socket;
- one slot cannot accept on another slot's listener;
- closing all sockets belonging to a slot works;
- one native process cannot close another process's socket;
- the owning native process can close it.
The test stubs net_tcp_xmit; it does not validate handshake, checksum, sequence, retransmission, or TCP parser correctness.
End-to-end evidence¶
The makefile forwards host TCP port 7007 to guest port 7 and host port 9016 to guest port 9016.
This provides integration paths for echo and CFS1 transfer.
Applications such as the browser, SSH banner experiment, and TLS/X25519 experiment also call the generic socket connect/send API.
Their existence demonstrates consumers of the API, not standards-complete TCP interoperability.
Concurrency¶
All TCP paths share the global packet builder g_tx.
The built-in echo state, socket table, transfer connection, and TX device buffer also have no general cross-CPU locking.
The current system relies on effectively serialized polling and syscall behavior.
Concurrent AP transmit or state mutation is not generally safe.
Recommended hardening¶
The immediate priorities are:
- validate TCP checksum on receive;
- centralize IPv4/TCP physical-length checks before any service dispatch;
- reject or correctly handle fragments before TCP;
- validate ACK numbers during handshake;
- require expected sequence numbers and implement reassembly;
- honor receive-window capacity instead of acknowledging discarded bytes;
- implement retransmission for data and FIN;
- implement peer FIN/RST and real close states;
- align maximum TCP payload with the 1514-byte Ethernet frame limit;
- unify echo, sockets, and transfer behind one TCP connection engine;
- add deterministic packet tests and malformed-input fuzzing.
Current limitations¶
The current TCP implementation is suitable for controlled ChrisOS/QEMU experiments, not hostile or lossy networks.
It lacks receive checksum validation, complete handshake validation, ordered stream reassembly, general retransmission, adaptive timeout, flow control, congestion control, options, robust teardown, RST handling, PMTU integration, and SMP-safe state ownership.
The project does have a real checksum generator, working happy-path SYN/data exchange, socket ownership controls, and a useful TCP-based file-transfer workflow.
Those capabilities should be stated precisely without implying a complete production TCP stack.
Revision note¶
This chapter was created against ChrisOS revision e05a17fd76333114a3fb5c2452f38ca747d4ac56. It documents the three distinct TCP state machines, their shared packet builder, the CFS1 transfer path, and the exact reliability and bounds gaps present in the inspected source.