The Network Times: 2026

Monday, 2 March 2026

Packet trimming Deep Dive - Part IV

Receive Network Processing Unit (Rx NPU)

Figure 9-4 illustrates a simplified receive-side processing pipeline, starting from the moment a Packet Header Vector (PHV), constructed by the Rx IFG, is delivered to the Receive Network Processing Unit (Rx NPU).

When the PHV arrives at the Rx NPU, it is dispatched to one of the Run-to-Completion (RTC) cores in the Packet Processing Array (PPA). Each RTC core processes the packet within a single execution context, allowing parsing, classification, lookup, and queuing decisions to be resolved without intermediate handoffs between processing stages.

The first task of the RTC parser is to perform deep inspection of the packet headers. While the Rx IFG has already extracted basic Layer-2 and Layer-3 information, the RTC parser determines whether the packet is tunneled and whether the switch itself is the tunnel termination point. To demonstrate this behavior, consider a VXLAN-encapsulated packet. The outer Ethernet and IP headers are used to forward the packet through the underlay network. If the outer destination IP address matches one of the local switch IP addresses, the device identifies itself as the tunnel endpoint. The tunneling protocol is recognized by examining the UDP header, where destination port 4789 indicates VXLAN. After the tunneling mechanism is identified, the outer headers are logically removed, and processing continues using the inner Ethernet and IP headers. These inner headers then form the basis for forwarding decisions. In this example, tunneling illustrates a scenario in which the switch operates as a Virtual Tunnel Endpoint (VTEP) in a multitenant scale-out backend network.

In parallel with deep parsing, traffic classification takes place. The packet is assigned an Internal Traffic Class (ITC) by the pre-classification process in the Rx IFG pipeline. The ITC is used solely for internal prioritization within the Rx NPU, such as memory access arbitration and scheduling of processing resources inside the pipeline. It influences how the packet progresses through the internal stages of the NPU but does not determine where the packet is buffered for transmission. The ITC is carried as metadata within the Packet Header Vector (PHV), ensuring that every internal bus and memory controller treats the packet according to its pre-assigned urgency as it traverses the NPU.

The Rx NPU pipeline, in turn, matches the DSCP field against the configured QoS classification policy. The result of this policy evaluation determines the Virtual Output Queue (VOQ) into which the packet will be enqueued. As described in the VOQ chapter, VOQs are organized by Traffic Class and destination, ensuring that congestion affecting one egress interface does not introduce head-of-line blocking for traffic destined to another. This separation decouples ingress buffering from egress congestion and preserves fairness under load.

At the same time, the RTC core performs a forwarding lookup against the Forwarding Information Base (FIB). The lookup resolves the egress interface and any associated forwarding attributes. Once the egress interface is known, the previously selected VOQ is mapped to the corresponding Output Queue (OQ) on that interface. This mapping follows the Traffic Class–to–egress priority relationship described earlier, ensuring that packets maintain consistent priority semantics from ingress classification through egress scheduling.

After the VOQ-to-OQ mapping is established, the Traffic Manager initiates a credit request toward the Tx NPU scheduler. This interaction follows the credit-based flow control model described in the VOQ chapter. Conceptually, the request indicates that a packet of a given size, associated with a specific egress port and priority level, is ready for transmission. The scheduler evaluates whether the egress port’s microscopic FIFO has sufficient available buffer space and whether any higher-priority packets are waiting to be transmitted. If the conditions allow, credits equal to the packet length are granted.

Only after credits are granted is the packet permitted to move from Unified Shared Memory (USM) toward the egress pipeline. This strict separation between enqueueing and transmission prevents buffer overcommitment and enforces priority ordering during congestion.

Once the Tx NPU scheduler grants the necessary credits, the packet is dequeued from the Unified Shared Memory and enters the Transmit Network Processing Unit (Tx NPU). Similar to the receive side, the transmit pipeline utilizes a Run-to-Completion (RTC) model within its own Packet Processing Array (PPA). This ensures that the final packet transformations are performed with the same deterministic, single-context efficiency as the initial ingress processing.

Upon entering the Tx NPU, the packet is dispatched to a Tx RTC core. The core's primary responsibility is header reconstruction and encapsulation. While the Rx NPU made the forwarding decision, the Tx NPU executes the "physical" rewrite. For a packet exiting a VTEP, this is where the RTC engine pushes the appropriate VXLAN, UDP, IP, and Ethernet headers onto the inner payload. Because this is a programmable RTC environment, the device can support complex, multi-label stacks, such as SRv6 or deep MPLS label impositions, without the "recirculation" penalties found in fixed-pipeline ASICs.

In addition to encapsulation, the Tx NPU performs a final round of Egress Policy Enforcement. This includes applying egress ACLs, updating packet counters for billing or monitoring, and inserting In-band Network Telemetry (INT) metadata if configured. This allows the switch to timestamp the packet at the precise moment of departure, providing nanosecond-accurate latency data.

The final stage of the Tx NPU involves the Output Queue (OQ) Scheduler. Even though the packet has already been "credited" for transmission, this local scheduler manages the final arbitration between different traffic classes sharing the same physical port. It ensures that a burst of low-priority bulk data does not jitter a high-priority data stream at the very last microsecond of the journey.

Finally, the fully formed packet is handed off to the MAC and PCS (Physical Coding Sublayer). Here, the digital data is serialized and mapped into PAM4 (Pulse Amplitude Modulation 4-level) symbols. These symbols are then modulated onto the physical medium, whether as electrical signals over a backplane or light pulses through an optical transceiver, completing the packet's journey through the Silicon One architecture.

In summary, the Rx NPU integrates tunnel awareness, forwarding lookup, QoS-based queuing, and credit-controlled admission into the egress pipeline within a single run-to-completion processing model. Internal Traffic Class governs how the packet is processed inside the NPU, while the QoS policy determines where the packet waits for transmission. This separation of responsibilities enables deterministic performance, scalable queuing, and strict priority enforcement across the switching fabric.

Figure 9-4: Rx NPU Pipeline.

Friday, 27 February 2026

Packet Trimming Deep Dive - Part III

Virtual Output Queue (VOQ)

The Silicon One VOQ Architecture

Instead of using dedicated deep interface buffers for packet queuing, Cisco Silicon One utilizes a Centralized Shared Memory architecture paired with a logical Virtual Output Queue (VOQ) mechanism. Because the VOQ concept is implemented within the Ingress (Rx) NPU entity, this queuing stage occurs after the initial ingress lookups but before the packet is switched across the internal fabric to the egress.

The VOQ model turns the traditional egress queuing model, where packets wait for serialization in a hardware buffer on the specific egress interface, upside down. While a VOQ is physically located on the ingress NPU, its ability to send traffic is controlled by the state of a small hardware Output Queue (OQ) on the egress interface.

Priority Mapping and Default State

As shown in Figure 9-3, a QoS policy can be created where a packet received on interface gi1/0/1 is assigned to Traffic Class 6 if the DSCP bits are set to EF (Expedited Forwarding). This configuration instantiates a VOQ specifically for that traffic class. In this hierarchy:

TC 7 (Control Plane/CS6): Mapped to OQ 1, the highest Strict Priority (Level 1).

TC 6 (DSCP-TRIMMED/EF): Mapped to OQ 2, the second-highest priority (Level 2).

By default, Silicon One enables VOQ 7 (Network Control) and VOQ 1 (Default/Best Effort). This ensures that critical control-plane traffic (DSCP 48/CS6) is guaranteed a high-priority path, to keep the network stable, while all unclassified data flows through the default VOQ. VOQs 5 – 2 are disabled by default.

Congestion Management and HoL Prevention

The VOQ mechanism is critical for handling "many-to-one" traffic patterns. For example, if four 800G ingress interfaces all send "elephant flows" to a single 800G egress interface, the egress OQ will become congested. Through a credit-based flow-control system, the egress port stops issuing credits to the specific ingress VOQs targeting it.

Because these queues are "virtualized" per output, this congestion does not cause Head-of-Line (HoL) blocking. Traffic destined for other, non-congested egress ports continues to receive credits and flow freely, even though they share the same physical ingress NPU.

Internal Isolation

The VOQ-OQ mechanism is entirely internal to the switch. Peer devices have no visibility into these internal queues; they only see the resulting serialized traffic stream and the DSCP/CoS markings in the packet headers.

Figure 9-3: Virtual Output Queue.

Wednesday, 25 February 2026

Packet Trimming Deep Dive - Part II

Receive Interface Group (Rx IFG)

Ingress Pre-Processing and Integrity

The Receive Interface Group (Rx IFG) is the ingress pre-processing stage that handles the incoming Ethernet bitstream before the packet enters the Packet Processing Array (PPA) of the Receive Network Processing Unit (Rx NPU) in the Cisco Silicon One architecture.

Processing begins at the Rx MAC. The Rx MAC reconstructs (“delimits”) the Ethernet frame from the Physical Coding Sublayer (PCS) bitstream and verifies frame integrity by computing a Frame Check Sequence (FCS) using the CRC-32 algorithm. If the computed FCS does not match the received FCS value, the frame is considered corrupted and is dropped immediately at ingress. If the CRC check succeeds, the frame is admitted for further processing.

Shallow classification and Traffic Class mapping

After frame validation, the Rx IFG identifies the Ethernet MAC header and detects the presence of IEEE 802.1Q VLAN tags. The Rx IFG performs shallow classification to efficiently manage hardware resources before deeper protocol parsing and forwarding decisions are executed in the Rx NPU. When an IEEE 802.1Q VLAN tag is present, the Rx IFG extracts the Priority Code Point (PCP) bits from the VLAN tag and maps them to an Internal Traffic Class (ITC). Based on this mapping, the frame is placed into a specific port-based local hardware FIFO queue.

If the frame does not carry a VLAN tag, or if no usable CoS information is available, the frame is assigned to a default Internal Traffic Class and buffered in a standard port-based FIFO queue.

At this stage, no forwarding lookup is performed; the purpose is limited to frame validation and shallow classification.

Port-based FIFO queues and packet buffering

The small port-based FIFO queues serve two purposes.

First, they provide prioritization for writing packet data into shared SRAM. Prioritization is required because a large number of packets, originating from many ingress ports and representing thousands of concurrent flows, may arrive simultaneously at the Rx IFG. The FIFO queues regulate this contention and determine the order in which packets are admitted into shared memory based on their assigned Internal Traffic Class.

Second, in parallel with the memory write operation, the Rx IFG generates a compact packet metadata structure known as the Packet Header Vector (PHV).

Packet Header Vector (PHV) Creation

The Packet Header Vector summarizes essential information about the packet without requiring full packet inspection. It is created in the Rx IFG to offload basic packet characterization from the Rx NPU parser, thereby reducing processing overhead in the programmable pipeline and enabling higher sustained packet rates. The PHV includes metadata such as:

Ingress interface and timestamp, and Packet length
Pointer to the packet’s cell chain in memory
Internal Traffic Class (ITC), VLAN ID, and EtherType
Router MAC hit indication

The Rx IFG operates as a pattern-matching engine rather than performing linear, sequential parsing of the entire packet. For example, when the EtherType field indicates IPv4 (0x0800), the Rx IFG recognizes that the IP Protocol field, used to identify the transport protocol such as TCP or UDP, is located at a fixed offset within the IPv4 header. It can therefore extract that field directly without scanning the header byte by byte.

Similarly, when an IEEE 802.1Q VLAN tag is present, the Rx IFG accounts for the additional header field and adjusts the parsing offset accordingly to locate the correct EtherType position before proceeding with further field extraction.

The Router MAC field in the PHV is a hit indicator that signals whether the destination MAC address matches one of the switch’s configured router MAC addresses. A positive match indicates that the packet is addressed to the switch itself and allows the Rx NPU to immediately invoke Layer 3 processing logic, bypassing Layer 2 forwarding paths.

Handoff to the Rx NPU

Once the PHV is constructed, it is dispatched to one of the Run-to-Completion (RTC) cores in the Packet Processing Array (PPA) of the Rx NPU. Because the Rx IFG includes a Pre-Parser, it can perform flow-based hashing. It examines the extracted L2/L3 header fields, computes a hash value, and ensures that all packets belonging to the same flow are directed to the same PPA core.

This preserves packet ordering while distributing the overall processing load across the chip. In Silicon One, this mechanism is commonly referred to as Local Target ID (LTID) generation. The IFG assigns a target ID to each packet, instructing the hardware which processing “lane” the packet should follow

Figure 9-2: Receive Interface Group Processes.

Friday, 20 February 2026

Packet Trimming Deep Dive - Part I

Introduction

The previous chapter introduced the Ultra Ethernet (UE) Transport Layer and its endpoint-centric congestion control mechanisms: Network Signaled Congestion Control (NSCC) and Receiver Credit-based Congestion Control (RCCC). This chapter moves down to the UE Network Layer and introduces Packet Trimming (PT).

While node-based approaches rely on NIC-to-NIC feedback loops, Packet Trimming allows network switches to actively intervene during periods of high utilization. Instead of silently dropping packets under congestion, the network provides an explicit and fast signal that enables immediate recovery.

The primary goal of Packet Trimming is to prevent incast congestion, a situation in which multiple ingress ports simultaneously overwhelm a single egress port. In AI and HPC workloads, many-to-one traffic patterns are common—for example, when multiple workers send data to a single parameter server. Under these conditions, egress buffers can be exhausted very quickly. In a best-effort network, this typically results in tail drops. The receiver then waits for a retransmission timeout, which introduces long tail latency and disrupts synchronization across distributed workloads. Packet Trimming replaces this silent packet loss with an explicit congestion signal that travels faster than the data itself.

The process begins at the source UE node. The NIC marks outgoing data packets with a DSCP-TRIMMABLE codepoint, indicating that the packet payload may be truncated if congestion occurs in the network. When such a packet enters a switch, it is initially treated as normal data traffic, for example classified into a low-priority traffic class and forwarded according to standard scheduling rules.

As the switch resolves the egress port, its scheduling logic continuously monitors congestion on that port. If a predefined congestion threshold is exceeded, the switch performs a precise operation. Instead of dropping the packet entirely, it discards only the payload while preserving the essential transport and network headers. For IPv4 traffic, the packet is typically truncated to 64 bytes, and for IPv6 traffic to 128 bytes. These sizes are chosen to ensure that all transport-level identification fields remain intact.

The Packet Delivery Sublayer (PDS) header carries the transport identity of the packet, including the Packet Sequence Number and the Packet Delivery Context identifiers. This information is sufficient for the receiver to detect packet loss and request retransmission. The switch does not need to preserve application semantics to signal congestion; it only needs to preserve transport-level identity.

Packet Trimming is effective only if the trimmed packet retains all information required for transport-level recovery. For this reason, Ultra Ethernet defines a minimum trim size that ensures the complete PDS header is preserved. This guarantees that sequence numbers and context identifiers are not lost during trimming.

In networks that use tunneling or encapsulation, such as VXLAN, additional headers must also be preserved. The receiver relies on this information to correctly demultiplex traffic and associate it with the appropriate transport context. As a result, the minimum trim size is a network-wide configuration parameter that depends on the transport protocol, IP version, and encapsulation methods used in the fabric. All switches in the network must apply the same value to ensure consistent behavior.

From a practical standpoint, trimming is most effective for large data packets. When a multi-kilobyte packet is reduced to a compact header-only frame, the data rate is reduced by orders of magnitude, while loss detection still happens within a single round-trip time. For small packets, the relative benefit of trimming is limited because the size reduction is much smaller.

After trimming the payload, the switch rewrites the DSCP field in the IP header from DSCP-TRIMMABLE to DSCP-TRIMMED. This rewrite explicitly signals that the packet has been trimmed and now represents congestion feedback rather than application data. The packet is then reclassified into a higher-priority traffic class before transmission.

This priority promotion is essential. Trimmed packets are forwarded using a medium- or high-priority queue so that they bypass the congestion that caused the trimming in the first place. Because trimmed packets are very small, prioritizing them does not increase congestion. Instead, it shortens the feedback loop by ensuring that congestion information reaches the destination with minimal delay.

When the trimmed packet arrives at the destination UE node, the transport layer immediately recognizes the DSCP marking and the absence of the payload. Using the preserved Packet Sequence Number from the PDS header, the receiver generates a selective negative acknowledgment with the retransmit flag set. This allows the sender to begin retransmission almost immediately, without waiting for a timeout.

By combining payload removal, DSCP rewrite, and priority forwarding, Packet Trimming enables fast, hardware-assisted recovery while maintaining high throughput. This capability is critical for large-scale AI training and HPC workloads, where low latency and tight synchronization directly impact performance.

The following sections describe how Packet Trimming and signal processing are implemented within the Cisco Silicon One G200 ASIC.

Optical to Digital Signal Processing

Figure 9-1 provides a high-level overview of the components involved in translating a received signal into an Ethernet bitstream. It traces the path of a signal encoded using Pulse Amplitude Modulation 4 (PAM4) as it travels from the optical domain, through electrical processing inside the ASIC, and finally to the RX MAC, where Ethernet frame boundaries are identified and the Cyclic Redundancy Check (CRC) is validated.

In Figure 9-1, an 800 Gbps optical transceiver is attached to front-panel interface E0, which belongs to Interface Group 0 (IFG0) covering interfaces E0–E3. Internally, each 800G interface is supported by eight 112 Gb/s SerDes lanes. These lanes are processed independently and later aggregated by the Physical Coding Sublayer (PCS) before reaching the RX MAC.

Phase 1 - Optical Reception: The transceiver receives an optical waveform that is already PAM4-modulated by the transmitting device. PAM4 encodes two bits per symbol by using four distinct optical symbol levels, allowing higher data rates without increasing the symbol rate.

Phase 2 - Photodiode Conversion: A photodiode converts the incoming optical signal into an analog electrical waveform, preserving the relative PAM4 symbol levels. Due to attenuation, dispersion, and noise introduced during transmission over fiber, this waveform may be distorted when it arrives at the receiver.

Phase 3 - Transceiver DSP Conditioning: To ensure the signal can reliably traverse the switch’s internal circuit board, the transceiver includes a small Digital Signal Processor (DSP). This DSP performs signal conditioning functions such as amplification, equalization, and retiming. The result is a cleaned PAM4 electrical signal, which is transmitted toward the ASIC across eight high-speed differential electrical lanes.

Phase 4 - ADC and SerDes Processing: Inside the G200 ASIC, each incoming electrical lane from the transceiver terminates at a dedicated SerDes RX slice, which forms the physical-layer front end of the ASIC. An Analog-to-Digital Converter (ADC) samples the incoming analog waveform, and SerDes logic performs equalization, clock recovery, and PAM4 symbol decoding to reconstruct the digital bitstream for that lane. At this stage, each SerDes lane produces a high-speed stream of digital data. To make this data manageable for internal logic, the SerDes widens the data path by converting the serial stream into parallel data. For example, instead of processing a single bit at a 112 Gb/s rate, the data may be represented as a 64-bit-wide word operating at approximately 1.75 GHz. This example illustrates the principle rather than a fixed architectural constant.

Phase 5 - PCS Lane Bonding: Because the full 800 Gbps signal is distributed across eight independent lanes, the Physical Coding Sublayer (PCS), alongside the RS-FEC sublayer, performs lane alignment, and deskew. The PCS then bonds the lane data together, producing a single continuous logical bitstream that represents the Ethernet interface. This aggregated bitstream is delivered to the RX MAC, which detects the Ethernet preamble and Start-of-Frame Delimiter (SFD), validates the Frame Check Sequence (CRC), and prepares the frame for further processing within the switching pipeline.

Why This Structure Is Necessary

No practical silicon device can process a serial data stream at hundreds of gigabits per second directly. Such operation would exceed thermal, power, and signal integrity limits.

By converting speed into width, the G200 trades extremely fast serial signaling for wide, lower-frequency parallel processing. In the example above, widening the data path to 512 bits (8 lanes × 64 bits) allows the ASIC to sustain the line-rate throughput while operating internal logic at a manageable clock rate. This architectural approach enables high performance without compromising power efficiency or reliability.

In modern high-speed Ethernet ASICs, the SerDes and PCS functions form the physical-layer front end of the chip, while the RX MAC operates only after the electrical lanes have been recovered, aligned, and bonded into a logical bitstream.

Figure 9-1: Optical to Digital Signal Processing.

Thursday, 5 February 2026

Ultra Ethernet: Receiver Credit-based Congestion Control (RCCC)

Introduction

Receiver Credit-Based Congestion Control (RCCC) is a cornerstone of the Ultra Ethernet transport architecture, specifically designed to eliminate incast congestion. Incast occurs at the last-hop switch when the aggregate data rate from multiple senders exceeds the egress interface capacity of the target’s link. This mismatch leads to rapid buffer exhaustion on the outgoing interface, resulting in packet drops and severe performance degradation.

The RCCC Mechanism

Figure 8-1 illustrates the operational flow of the RCCC algorithm. In a standard scenario without credit limits, source Rank 0 and Rank 1 might attempt to transmit at their full 100G line rates simultaneously. If the backbone fabric consists of 400G inter-switch links, the core utilization remains a comfortable 50% (200G total traffic). However, because the target host link is only 100G, the last-hop switch (Leaf 1B-1) becomes an immediate bottleneck. The switch is forced to queue packets that cannot be forwarded at the 100G egress rate, eventually triggering incast congestion and buffer overflows.

While "incast" occurs at the egress interface and can resemble head-of-line blocking, it is fundamentally a "fan-in" problem where multiple sources converge on a single receiver. Under RCCC, standard Explicit Congestion Notification (ECN) on the last-hop switch's egress interface is typically disabled for this traffic class. The reasoning is twofold:

Redundancy: In Ultra Ethernet, ECN is the primary signal for NSCC to adjust the Congestion Window (CWND) and rotate the Entropy Value (EV) to trigger packet-level load balancing across the fabric.

Path Convergence: At the last-hop switch, rotating the EV is ineffective because there is only a single physical path to the destination. Since RCCC provides a more granular, proactive mechanism to throttle senders based on the receiver's actual capacity, the reactive "slow down" signaling of ECN becomes unnecessary at this stage. By disabling ECN here, the receiver (Target) takes full responsibility for flow management, ensuring that the fabric remains clear of congestion markers that might otherwise trigger unnecessary path hunting.

Credit Allocation and Flow

Instead of relying on late-stage ECN signaling, the RCCC algorithm proactively throttles senders by granting credits that match the physical transport speed of the target's connection.

Discovery: When Rank 2 receives data, it identifies the sources via the CCC_ID field in the RUD_CC_REQ (the specific request type used when RCCC is enabled) and adds them to its Active Sender Table.

Calculation: The algorithm divides the total available bandwidth, for a 100Gbps link, this is roughly 12.5 GB/, among the active senders. In this example, each sender is allocated 6.25 GB/s (50Gbps) worth of credits.

Granting: These credits are transmitted back to the sources via ACK_CC packets once data is successfully committed to Rank 2’s memory.

Enforcement: Upon receiving the ACK_CC, the Congestion Control Context (CCC) associated with the sender’s Packet Delivery Control (PDC) updates its local credit table. The PDC only permits transmission based on these available credits, effectively capping the individual sender's rate at 50G. This ensures that when combined with the other sender, the aggregate rate at the receiver does not exceed its 100G link capacity.

This credit-grant loop is continuous. The RUD_CC_REQ carries "backlog" information, telling the target exactly how much data is waiting in the source's queue. By dynamically adjusting grants based on this feedback, RCCC ensures the backend network remains lossless.

Figure 8-1: RCCC: Destination Flow Control.

Source RCCC Operation

The RCCC operation from the perspective of source UET Node-A begins when an application on Rank 0 initiates a 256 MB Remote Memory Access (RMA) write operation toward Rank 2. This request is handled by the Semantic Sublayer (SES), which translates the high-level command into ses_pds_tx_req for the Packet Delivery Sublayer (PDS). In our example, the PDS Manager determines that no communication channel currently exists between the Fabric Endpoints used for these connections, so it allocates a new Packet Delivery Control (PDC) from its general pool with the PDC identifier 0x4001. Simultaneously, it requests a Congestion Control Context (CCC) from the Congestion Management System (CMS), resulting in a dedicated context, CCC_ID = 0xA1, being configured and bound to the new PDC.

Once PDC and CCC are established, the system tracks the pending data through a two-tier backlog system. In our example, PDC 0x4001 updates its delta backlog with the full 256 MB of the request, which is then added to the CCC’s global backlog. This global value represents the total volume of data currently waiting for transport across all PDCs managed by that specific context. Because this is the start of the transaction, the global backlog moves from zero to 256 MB, establishing the total "demand" the source is prepared to place on the network.

In our example, new contexts are pre-provisioned with initial credits scaled to the Bandwidth-Delay Product (BDP). While the theoretical capacity of a 100G link is 12.5 GB/s, the initial "pipe-cleaning" burst is much smaller, specifically 12.5 KB in this scenario. This value represents a safe, conservative fraction of the total BDP, ensuring that the source can trigger the feedback loop without the risk of overwhelming the receiver's buffers or the last-hop switch before the control loop fully engages. The CCC authorizes PDC 0x4001 to transmit this initial amount, subtracts it from the current cumulative credits, and updates the global backlog to show that this small portion is now in-flight, leaving 255.987.500 bytes remaining in the queue.

With this authorization from the CCC, the PDC passes the work request to the NIC, which fetches the data from memory and prepares the packet for transmission. In our example, the FEP Fabric Addresses (FA) are encoded into the IP header’s source and destination IP address fields, and the DSCP bits are configured to correspond to the TC-LOW traffic class. Additionally, the ECN bits are set to reflect that the packet is ECN-capable, ensuring visibility for Network Signaled Congestion Control (NSCC) if needed. The type of the PDS request is set to RUD_CC_REQ, which requires a pds.req_cc_state field. In our example, this field carries the CCC_ID (0xA1) and the Credit Target, which describes the size of the backlog of the sender CCC. By including these parameters, the source explicitly informs the target of its total remaining data, allowing the receiver to calculate and return the next set of credit grants to keep the pipeline moving.

Note: Since the source does not yet have information regarding the PDC on the remote target, it populates the pdc_info field with a value of 0x0 for the Destination PDC ID (DPDCID) for notifying the target that its new PDC ID must be taken from the global PDC pool. Furthermore, the SYN bit remains set until the first ACK_CC message is received, signaling to the target that the connection handshake and credit-granting loop are in the initialization phase.

Figure 8-2: Source RCCC Processing.

Target RCCC Operation – PDS Request

When the initial packet arrives at destination Node-B, the PDS Manager first checks for an existing PDC associated with the incoming connection from Fabric Address (FA) 10.0.0.1 and SPDCID 0x4001. Because no such PDC exists, the PDS Manager identifies this as a new connection request. The value of 0x0 in the pdc_info field instructs the target to allocate a General type PDC, ensuring the local delivery control matches the source's PDC type.

Since no Congestion Control Context (CCC) currently exists for this specific FEP-to-FEP connection, the PDS requests the CMS to allocate a new one. The CMS assigns CCC_ID 0xB1 and creates an entry source CCC_ID-specific entry in the Active Sender Table. This entry describes the source address (FA 10.0.0.1), and the assigned traffic class (TC-LOW) from the IP header. Besides, entry for source CCC_ID 0xA1 described in PDS headers tells the source backlog size as credit_target with value 255.987.500.

Simultaneously, the NIC extracts the semantic information from the SES header to identify the required operation. In our example, it recognizes a UET_WRITE command and determines the target memory address for the incoming data. Once the packet payload is verified, the data is forwarded to the High-Bandwidth Memory (HBM) Controller, where it waits for its turn to be committed to the physical memory.

Figure 8-3: Target RCCC Processing – PDS Request.

Target RCCC Operation – Credit Assignment

After receiving confirmation from the SES regarding the completed memory operation, the PDS prepares the response using an ACK_CC message. The CMS must now determine how much data the source is permitted to send in its next burst. In our example, the CMS allocates 12.5 KB of credits for CCC_ID 0xA1.

The math behind this allocation is a function of the receiver’s total capacity and the time-granularity of the control loop. While the NIC provides a 100 Gbps (12.5 GB/s) "pipe," the receiver does not grant a full second of data at once, as doing so would bypass the congestion control mechanism. Instead, it grants data in "time-slices", in this scenario, representing 1 microsecond of transmission. By dividing the total bandwidth by the number of active senders for that specific time-slice, the receiver ensures that the aggregate "demand" never exceeds the physical capabilities of the link.

In our example, with only one active sender, the calculation is:

(12.5 GB/s x 0.000001 s) ÷ active senders = 12.5 KB

The RCCC algorithm is designed for dynamic fairness. Though not explicitly shown in Figure 8-4, if Rank 0 had a simultaneous transfer in progress, the Active Sender Table would list two sources. The CMS would then divide that same 1-microsecond "slice" between them, reducing the granted credit per source to 6.25 KB. This prevents "incast" congestion by ensuring that even if multiple sources transmit at once, their combined throughput matches exactly what the receiver can ingest.

The PDS defines this response by setting the pds.cc_type to CC_CREDIT. The pds.ack_cc_state field is populated with this calculated credit value, while the ooo_count field tracks any Out-of-Order packets. To ensure this information is not delayed by standard data traffic, the DSCP bits in the IP header are set to TC-High. This gives the ACK_CC message "express" priority across the backend fabric, minimizing the time the source spends waiting for new credits and maintaining a high-performance, steady-state flow.

Crucially, the Target populates its own local PDC ID (0x4011) into the Source PDC Identifier (SPDCID) field of the PDS Prologue header. By doing so, it provides the return address necessary for the source to transition out of its initial "discovery" state.

Figure 8-4: Target RCCC Processing – ACK_CC Message Reply.

Source-Side Processing of ACK_CC

When the ACK_CC message arrives at the source, the NIC identifies the target FEP based on the destination IP address. However, for high-speed internal processing, it uses the DPDCID in the PDS header as a local handle to jump directly to the correct PDC Context. From this entry, the NIC automatically resolves the CCC_ID associated with that specific PDC.

Once the correct CCC entry is identified, the source processes the new credit information. In our example, the receiver has sent a new Cumulative Credit value of 25,000 bytes. To determine the currently available window, the source performs a simple subtraction:

Incremental Credit = Received Cumulative Credit – Local Cumulative Credit

By subtracting the previously recorded 12,500 bytes from the new 25,000 bytes, the source identifies an incremental grant of 12,500 bytes. The CCC then authorizes the PDC to transmit this amount. Simultaneously, the Global Backlog is updated by subtracting these 12,500 bytes from the remaining 255,987,500 bytes, keeping the sender’s demand signal accurate for the next request.

The PDC informs the NIC that it is cleared to construct packets fitting this allowed credit size (respecting the NIC’s MTU). The NIC fetches the data from memory, packetizes it, and transports it to the destination. This control loop continues—updating demand and receiving cumulative grants—until the entire backlog has been transported and acknowledged.

Once the job is complete, the PDC context is closed. If no other PDCs are currently associated with that CCC_ID, the CCC is also closed. This hierarchical teardown ensures that no unnecessary hardware resources or bandwidth are reserved in the AI Fabric once the work is done.

Saturday, 31 January 2026

Ultra Ethernet: NSCC Destination Flow Control

Figure 6-14 depicts a demonstrative event where Rank 4 receives seven simultaneous flows (1). As these flows are processed by their respective PDCs and handed over to the Semantic Sublayer (2), the High-Bandwidth Memory (HBM) Controller becomes congested. Because HBM must arbitrate multiple fi_write RMA operations requiring concurrent memory bank access and state updates, the incoming packet rate quickly exceeds HBM’s transactional retirement rate.

This causes internal buffers at the memory interface to fill, creating a local congestion event (3). To prevent buffer overflow, which would lead to dropped packets and expensive RMA retries, the receiver utilizes NSCC to move the queuing "pain" back to the source. This is achieved by using pds.rcv_cwnd_pend parameter of the ACK_CC header (4). The parameter operates on a scale of 0 to 127; while zero is ignored, a value of 127 triggers the maximum possible rate decrement. In this scenario, a value of 64 is utilized, resulting in a 50% penalty relative to the newly acknowledged data.

Rather than directly computing a new transport rate, the mechanism utilizes a three-phase process to define a restricted Congestion Window (CWND). This reduction in CWND inherently forces the source to drain its inflight bucket to maintain protocol compliance and synchronize the injection rate with the HBM's processing capacity. The process begins by calculating the newly_rcvd_bytes, representing the data volume acknowledged by the incoming ACK_CC. This is the delta between the rcvd_bytes of the predecessor ACK_CC (12,288 bytes) and the newest rcvd_bytes (16,384 bytes), totaling 4,096 bytes (A).

In the next phase, the logic multiplies 4,096 bytes by the rcv_cwnd_pend value of 64, resulting in a product of 262,144. Applying a bit-shift of 7 (equivalent to dividing by 128) yields a penalty of 2,048 bytes (B). This penalty is then subtracted from the current CWND of 75,776, establishing a new, throttled CWND of 73,728 bytes (C).

In a stable state, the CWND and the inflight bucket are typically equal in size; consequently, immediately following the decrement, the current inflight bucket exceeds the newly defined CWND limit by 2,048 bytes. This state violates the fundamental transport rule where the CCC allows the PDC to transmit data only when the inflight bucket is less than or equal to the CWND (5). In response, the PDC must suspend transmission, waiting for the destination to acknowledge enough packets to reduce the inflight bucket size to be less than or equal to size of the new CWND (6).

This pause allows the HBM controller the necessary time to clear its transaction queue. Only once the inflight level has drained to meet the new CWND ceiling can the CCC authorize the PDS to resume data transport. The rc-flag (Restore CWND) when set, it signals that after flow congestion control event, the original CWND can be utilized again.

Figure 6-14: NSCC: Destination Flow Control.

NSCC Mechanism Summary

The Network-Signaled Congestion Control framework ensures high-performance data transfer by balancing the real-time Inflight Load against a dynamic Congestion Window (CWND). By utilizing proactive feedback from the fabric and the destination, the system maintains line-rate performance while preventing buffer overflow and high tail latency.

Proportional and Fast Increase: These methods are utilized when the network is underloaded, characterized by a lack of ECN-CE signals and queuing delays below the target threshold. Proportional Increase scales the CWND based on the gap between measured and target delays to optimize utilization. Fast Increase employs exponential growth to quickly reclaim bandwidth when the network remains significantly underutilized for a duration.

Fair Increase: This method is initiated as congestion subsides to ensure an equitable recovery among competing flows. By adding a fixed, constant amount to the CWND of every active flow, it allows flows with smaller windows to grow at a faster relative rate, eventually leading all participants to converge on a fair share of the available bandwidth.

Multiplicative Decrease: This action is used to protect the fabric during periods of high pressure, specifically when queuing delay exceeds targets and ECN feedback indicates stagnant queues. It slashes the CWND proportionally to the measured buffer excess, rapidly shedding load to return the network queue to its target occupancy level within a single Round-Trip Time.

Destination Flow Control (NSCC Receiver Penalty): This mechanism addresses bottlenecks at the receiver’s hardware level, such as the High-Bandwidth Memory (HBM) controller. By applying a penalty via the rcv_cwnd_pend parameter, the receiver forces the source to reduce its CWND based on a percentage of the newly acknowledged data. This pauses new data injections until the destination's transaction queues have drained, moving the queuing pressure from the memory controller back to the source.

CWND Restoration: The Restore CWND method, triggered by the rc-flag, allows a flow to immediately resume its original transmission rate once a congestion event has passed. This prevents the flow from having to slowly ramp back up through increase phases, ensuring that the system returns to peak efficiency as soon as the bottleneck, whether in the fabric or at the destination, is resolved.

Thursday, 29 January 2026

Ultra Ethernet: Inflight Bytes and CWND Adjustment

Inflight Packet Adjustment

Figure 6-12 depicts the ACK_CC header structure and fields. When NSCC is enabled in the UET node, the PDS must use the pds.type ACK_CC in the prologue header, which serves as the common header structure for all PDS messages. Within the actual PDS ACK_CC header, the pds.cc_type must be set to CC_NSCC. The pds.ack_cc_state field describes the values and states for service_time, rc (restore congestion CWND), rcv_cwnd_pend, and received_bytes. The source specifically utilizes the received_bytes parameter to calculate the updated state for inflight packets.

The CCC computes the reduction in the inflight state by subtracting the rcvd_bytes value received in previous ACK_CC messages from the rcvd_bytes value carried within the latest ACK_CC message. As illustrated in Figure 6-12, the inflight state is decreased by 4,096 bytes, which is the delta between 16,384 and 12,288 bytes.

Recap: In order to transport data to network, the Inflight bytes must be less than CWND size.

Figure 6-12: NSCC: Inflight Bytes adjustment.

CWND Adjustment

A single, shared Congestion Window (CWND) regulates the total volume of bytes across all PDCs that are permitted for transmission to the backend network. The transport rate and network performance are continuously monitored and serve as the baseline information for dynamic CWND adjustments. The primary objective of these adjustments is to maintain minimal queue depth to eliminate queuing delay while ensuring the backend network is not overloaded, thereby guaranteeing lossless, line-rate packet transport.

Figure 6-13 illustrates the ACK_CC message structure, specifically the pds.flags field, which indicates whether a switch in the backend network has experienced congestion in an outgoing interface queue. The m-flag is set when the packet being acknowledged carries ECN-CE bits within the ToS field of the IP header.

In addition to the ECN state, the CWND adjustment algorithm compares the measured Queuing Delay against the Target Delay. As a recap from the Inflight section, the Queuing Delay is calculated as follows:

Queuing Delay = Ack_arrival_time - pkt_tx_state – service_time

The pkt_tx_state is recorded by the PDS and passed to the CCC, which derives the queuing delay by subtracting both the transmission timestamp and the service_time from the ACK arrival time.

The service_time is measured at the target by calculating the delta between the packet reception time (rx_time) and the response transmission time (tx_time). This value is encoded into the service_time parameter within the pds.ack_cc_state field. It represents the total processing overhead for the Packet Delivery and Semantic Sublayers to handle the RMA operation, including header extraction, memory addressing, data writing, SES response generation, and ACK_CC construction. By isolating this processing time, the source can accurately determine the true delay caused strictly by the network fabric.

Proportional and Fast Increase Logic

When the m-flag is clear (indicating no ECN-CE bits are detected) and the calculated Queuing Delay is less than the Target Delay, the network utilization is considered below its optimum. In this state, the CWND size increment is directly related to the magnitude of the difference between the measured queuing delay and the target delay.

A large difference between these two values indicates that network utilization is low and the path can handle a significantly higher volume of flows. The NSCC algorithm responds by increasing the CWND at a rate proportional to this gap, allowing more room for inflight packets. As the measured delay approaches the target delay and the difference narrows, the rate of increase automatically slows down to stabilize the flow.

In scenarios where the network remains significantly underloaded for a duration, such as when competing flows terminate, the system can escalate to a fast_increase. This mode employs exponential growth to quickly converge on the newly available bandwidth, remaining active until the system detects the first signs of incipient congestion.

Fair increase Logic

The fair_increase action is initiated when the system detects that a congestion event is subsiding. This state occurs when ECN signals indicate that the network queue has drained below the configured threshold, even if a previous packet experienced congestion. This mechanism is primarily designed to prevent the transmission rate from "undershooting" the actual network capacity during the recovery phase.

In this mode, the NSCC algorithm performs an additive increase rather than a proportional one. By adding a constant, fixed amount to the CWND, the system promotes fairness among competing flows. Because every flow receiving the same signal increases its CWND by the identical fixed value, flows with smaller windows experience a larger relative growth rate compared to those with larger windows. This ensures that all active flows eventually converge toward an equitable share of the available bandwidth.

Multiplicative Decrease Logic

The multiplicative_decrease action is triggered when the measured Queuing Delay exceeds the Target Delay and ECN feedback indicates that the network queue is not effectively decreasing. In this state, the average delay serves as a direct metric for the volume of excess data currently enqueued beyond the desired threshold.

The NSCC algorithm reacts by reducing the CWND proportionally to this measured queuing excess. By directly tying the window reduction to the specific magnitude of the buffer overflow, the system can rapidly shed load to alleviate congestion. When all competing flows execute this coordinated reduction, the objective is to clear the bottleneck and return the queue to its target occupancy level within approximately one Round-Trip Time (RTT).

Figure 6-13: NSCC: CWND Adjustment.

NSCC Summary

The Network-Signaled Congestion Control (NSCC) framework is designed to solve the fundamental challenge of high-performance networking: maximizing throughput while maintaining near-zero latency. This objective is achieved by managing a continuous balance between Inflight Load and the CWND Budget.

The system continuously adjusts transmission rates to ensure optimal network utilization. By proactively responding to fabric signals, NSCC maintains line-rate performance and prevents bufferbloat (a jam before buffer overflow), ensuring that packets spend as little time as possible waiting in switch queues. The core gatekeeper rule of the transport layer is defined by the relationship between two variables. Inflight bytes represent the real-time volume of data currently transiting the network, while the Congestion Window (CWND) represents the total data budget the network can safely handle. The source calculates the current Inflight state by tracking cumulative received_bytes from ACK_CC messages, and new data is only injected into the fabric when the Inflight count is lower than the CWND.

The sender dynamically scales the CWND budget using a sophisticated state machine based on the m-flag (ECN-CE) and measured Queuing Delay. It utilizes Proportional or Fast Increase to fill available bandwidth when the network is underutilized, ensuring the pipe remains full. As congestion clears, the Fair Increase mechanism ensures multiple flows converge to an equitable share of the pipe. Conversely, the Multiplicative Decrease action aggressively slashes the window when buffers overflow to protect the fabric from packet loss.

Tuesday, 27 January 2026

Ultra Ethernet: Network-Signaled Congestion Control (NSCC) - Overview

Network-Signaled Congestion Control (NSCC)

The Network-Signaled Congestion Control (NSCC) algorithm operates on the principle that the network fabric itself is the best source of truth regarding congestion. Rather than waiting for packet loss to occur, NSCC relies on proactive feedback from switches to adjust transmission rates in real time. The primary mechanism for this feedback is Explicit Congestion Notification (ECN) marking. When a switch interface's egress queue begins to build up, it employs a Random Early Detection (RED) logic to mark specific packets. Once the buffer’s Minimum Threshold is crossed, the switch begins randomly marking packets by setting the last two bits of the IP header’s Type of Service (ToS) field to the CE (11) state. If the congestion worsens and the Maximum Threshold is reached, every packet passing through that interface is marked, providing a clear and urgent signal to the endpoints.

The practical impact of this mechanism is best illustrated by a hash collision event, such as the one shown in Figure 6-10. In this scenario, multiple GPUs on the left-hand side of the fabric transmit data at line rate. Due to the specific entropy of these flows, the ECMP hashing algorithms on leaf switches 1A-1 and 1A-2 inadvertently select the same uplink to Spine 1A. Because all destination GPUs are concentrated on leaf switch 1B-1, the spine is forced to aggregate these incoming flows—totaling 500 Gbps—into a single outgoing interface. This bottleneck causes the queue to fill rapidly. Consequently, Spine 1A marks packets destined for Rank 9 and Rank 5 with ECN-CE. When these marked packets reach the receiver, the Packet Delivery Service (PDS) detects the congestion signal and reflects it back to the source by setting the pds.m flag in the acknowledgement (ACK) message.

The second signaling mechanism is based on measured queuing delay, which provides a granular view of fabric pressure even when ECN marks are not present. The algorithm calculates this by measuring the current Round-Trip Time (RTT) and subtracting the Base_RTT—the known minimum RTT of an uncongested path. This difference (Delta RTT) represents the time a packet spent sitting in switch buffers. By isolating the queuing delay from the total propagation time, the algorithm can detect the earliest stages of buffer buildup with high precision.

To manage these signals effectively, the algorithm maintains a constant record of the inflight packet state, tracking every byte transmitted to the network that has not yet been acknowledged or NACKed by the receiver. By synthesizing these three critical factors, ECN-CE signals, calculated queuing delay, and the volume of packets in flight, the NSCC algorithm dynamically adjusts the Congestion Window (CWND). This data allows the algorithm to decide precisely when a PDS is permitted to inject new data into the fabric and, if necessary, to rotate the Entropy Value (EV) to steer traffic toward underutilized paths, effectively resolving the collision and restoring optimal flow.

Figure 6-10: NSCC: Link Congestion due to Hash Collision.

The Overview of the NSCC Control Loop

Building on the previous overview, this section examines the granular mechanics of the NSCC process. Figure 6-11 illustrates the source-side operations as various Ranks initiate communication over the backend fabric. In this scenario, data from Ranks 0 and 8 is managed by Packet Delivery Context (PDC) 0x4001, Rank 2 is handled by PDC 0x4002, and Ranks 1 and 3 are assigned to PDC 0x4003.Each rank is tasked with transferring 4,096 KB of data. While abstracted in the diagram, the process begins when an application executes a fi_write RMA operation. This request is passed to the Semantic Sublayer (SES), which translates the intent into a UET_WRITE operation before handing it off to the PDC layer. Upon receiving new data, the PDC notifies the Congestion Control Context (CCC) Manager within the Congestion Management Sublayer (CMS) of a delta backlog (Steps 1a–c). This delta represents the volume of unsent data waiting in the PDC buffers that must be added to the total CCC backlog.

The CMS then acts as the gatekeeper; it compares the current inflight bytes against the Congestion Window (CWND). If the volume of data currently on the wire is less than the CWND, the CCC scheduler permits data transport (Step 2). In our example, there is sufficient headroom in the window, allowing the scheduler to authorize PDC 0x4001 to transmit. As the packet is dispatched, the hardware records the precise transmission time and injects the Entropy Value (EV) into the header to facilitate fabric load balancing (Phase 3). Simultaneously, the Inflight state is incremented and the backlog is decremented to reflect the data now transiting the network (Phases 4 and 5).The receiver processes the incoming packet and generates an ACK_CC message (Step 6). If the packet arrived with ECN-CE bits set by a switch, the receiver sets the pds.m flag in the ACK to signal that congestion was manifested. In this specific example, no congestion is encountered, so the pds.m bit remains unset. Crucially, the ACK_CC includes the service-time, the internal processing delay at the receiver—and a cumulative byte count to inform the source of the total data successfully received.

When the source receives the ACK_CC, it logs the arrival time (Step 7) and updates the CCC state. It decreases the inflight counter based on the rcvd_bytes value and adjusts the CWND. The adjustment is governed by two factors: the state of the ECN-CE bits and the measured Queuing Delay relative to the target delay. The Queuing Delay is calculated as:

Queuing Delay = Ack_arrival_time - pkt_tx_state – service_time

When packet trimming is used the default target delay is the same as configured base_rtt. Without packet trimming the target delay is base_rtt * 0.75. The CWND adjustment options are explained in the next section.

This autonomous, self-adjusting control loop represents a sophisticated implementation of Intent-Based Networking (IBN) at the transport layer. The high-level "intent" is simple: the reliable delivery of data between Ranks at line rate with minimal tail latency. To fulfill this, the NSCC algorithm operates as a real-time, closed-loop system—monitoring network feedback, analyzing fabric pressure, and adapting injection rates without human intervention. By offloading this decision-making to the Congestion Management Sublayer (CMS), the fabric becomes self-optimizing, ensuring that even in the face of unpredictable hash collisions, the network remains a transparent utility for the application.

Figure 6-11: NSCC Operation.

The following section concludes our exploration of NSCC by detailing the specific fields within the ACK_CC header and illustrating how the source-side state machine transitions between different congestion levels. While the overview provided here is sufficient to understand the fundamental operations of NSCC, the subsequent deep dive is intended for those who require bit-level architectural details.

While NSCC serves as the primary proactive mechanism for modulating flow at the source, it is only one part of the Ultra Ethernet "congestion toolbox." To ensure total fabric reliability, UEC employs additional layers of defense, such as Receiver Credit-based Congestion Control (RCCC) and Packet Trimming. These mechanisms are designed to handle specific scenarios where proactive rate-limiting isn't enough, providing the "emergency" recovery needed to maintain near-line-rate performance. Each of these solutions will be explored in detail in the upcoming chapters.

Tuesday, 13 January 2026

Ultra Ethernet: Congestion Control Context

Ultra Ethernet Transport (UET) uses a vendor-neutral, sender-specific congestion window–based congestion control mechanism together with flow-based, adjustable entropy-value (EV) load balancing to manage incast, outcast, local, link, and network congestion events. Congestion control in UET is implemented through coordinated sender-side and receiver-side functions to enforce end-to-end congestion control behavior.

On the sender side, UET relies on the Network-Signaled Congestion Control (NSCC) algorithm. Its main purpose is to regulate how quickly packets are transmitted by a Packet Delivery Context (PDC). The sender adapts its transmission window based on round-trip time (RTT) measurements and Explicit Congestion Notification (ECN) Congestion Experienced (CE) feedback conveyed through acknowledgments from the receiver.

On the receiver side, Receiver Credit-based Congestion Control (RCCC) limits incast pressure by issuing credits to senders. These credits define how much data a sender is permitted to transmit toward the receiver. The receiver also observes ECN-CE markings in incoming packets to detect path congestion. When congestion is detected, the receiver can instruct the sender to change the entropy value, allowing traffic to be steered away from congested paths.

Both sender-side and receiver-side mechanisms ultimately control congestion by limiting the amount of in-flight data, meaning data that has been sent but not yet acknowledged. In UET, this coordination is handled through a Congestion Control Context (CCC). The CCC maintains the congestion control state and determines the effective transmission window, thereby bounding the number of outstanding packets in the network. A single CCC may be associated with one or more PDCs communicating between the same Fabric Endpoint (FEP) within the same traffic class.

Initializing Congestion Control Context (CCC)

When the PDS Manager receives an RMA operation request from the SES layer, it first checks whether a suitable Packet Delivery Context (PDC) already exists for the JobID, destination FEP, traffic class, and delivery mode. If no matching PDC is found, the PDS Manager allocates a new one.

For the first PDC associated with a specific FEP-to-FEP flow, a Congestion Control Context (CCC) is required to manage end-to-end congestion. The PDS Manager requests this context from the CCC Manager within the Congestion Management Sublayer (CMS). Upon instantiation, the CCC initially enters the IDLE state, containing basic data structures without an active configuration.

The CCC Manager then initializes the context by calculating values and thresholds, such as the Initial Congestion Window (Initial CWND) and Maximum CWND (MaxWnd), using pre-defined configuration parameters. Once these initial source states for the NSCC are set, the CCC is bound to the corresponding PDC.

When fully configured, the CCC transitions to the READY state. This transition signals that the CCC is authorized to enforce congestion control policies and monitor traffic. The CCC serves as the central control structure for congestion management, hosting either sender-side (NSCC) or receiver-side (RCCC) algorithms. Because a CCC is unidirectional, it is instantiated independently on both the sender and the receiver.

Once in the READY state, the PDC is permitted to begin data transmission. The CCC maintains the active state required to regulate flow, enabling the NSCC and RCCC to enforce windows, credits, and path usage to prevent network congestion and optimize transport efficiency.

Note: In this model, the PDS Manager acts as the control-plane authority responsible for context management and coordination, while the PDC handles data-plane execution under the guidance of the CCC. Once the CCC is operational, RMA data transfers proceed directly via the PDC without further involvement from the PDS Manager.

Figure 6-6: Congestion Context: Initialization.

Calculating Initial CWND

Following the initialization of the Congestion Control Context (CCC) for a Packet Delivery Context (PDC), specific configuration parameters are used to establish the Initial Congestion Window (CWND) and the Maximum Congestion Window (MaxWnd).

The Congestion Window (CWND) defines the maximum number of "in-flight" bytes, data that has been transmitted but not yet acknowledged by the receiver. Effectively, the CWND regulates the volume of data allowed on the wire for a specific flow at any given time to prevent network saturation.

The primary element for computing the CWND is the Bandwidth-Delay Product (BDP). To determine the path-specific BDP, the algorithm selects the slowest link speed and multiplies it by the configured base Round-Trip Time (config_base_rtt):

BDP = min(sender.linkspeed, receiver.linkspeed) x config_base_rtt

The config_base_rtt represents the latency over the longest physical path under zero-load conditions. This value is a static constant derived from the cumulative sum of:

Serialization delays (time to put bits on the wire)
Propagation delays (speed of light through fiber)
Switching delays (internal switch traversal)
FEC (Forward Error Correction) delays

Setting MaxWnd

The MaxWnd serves as a definitive upper limit for the CWND that cannot be exceeded under any circumstances. It is typically derived by multiplying the calculated BDP by a factor of 1.5.While a CWND equal to 1.0 x BDP is theoretically sufficient to saturate a link, real-world variables, such as transient bursts, scheduling jitter, or variations in switch processing, can cause the link to go idle if the window is too restrictive. UET allows the CWND to grow up to 1.5 x BDP to maintain high utilization and accommodate acknowledgment (ACK) clocking dynamics.

Example Calculation: Consider a flow where the slowest link speed is 100 Gbps and the config_base_rtt is 6.0 µs.

Calculate BDP (Bits): 100 x 109 bps x 0.000006 s = 600,000 bits

Calculate BDP (Bytes): 600,000 / 8 = 75,000 bytes

Calculate MaxWnd: 75,000 x 1.5 = 112,500 bytes

Note on Incast Prevention: While the "ideal" initial CWND is 1.0 x BDP, UET allows the starting window to be configured to a significantly smaller value (e.g., 10–32 KB or a few MTUs). This configuration prevents Incast congestion, a phenomenon where the aggregate traffic from multiple ingress ports exceeds the physical capacity of an egress port. By starting with a conservative CWND, the system ensures that the switch's egress buffers are not exhausted during the first RTT, providing the NSCC algorithm sufficient time to measure RTT inflation and modulate the flow rates.

A common misconception is that the BDP limits the transmission rate. In reality, the BDP defines the volume of data required to keep the "pipe" full. While the Initial CWND may be only 75,000 bytes, it is replenished every RTT. At a 6.0 µs RTT, this volume translates to a full 100 Gbps line rate:

600,000 bits / 6.0 µs = 600,000 / 0.000006 = 100 × 109 bps = 100 Gbps

Therefore, a window of 1.0 x BDP achieves 100% utilization. The 1.5 x BDP (MaxWnd) simply provides the necessary headroom to prevent the link from going idle during minor acknowledgment delays.

Figure 6-7: CC Config Parameters, Initial CWND and MaxWnd.

Calculating New CWND

When the network is uncongested, indicated by a measured RTT remaining near the base_rtt, the NSCC algorithm performs an Additive Increase (AI) to grow the CWND. To ensure fairness across the entire fabric, the algorithm utilizes a universal Base_BDP parameter rather than the path-specific BDP.

The Base_BDP is a fixed protocol constant (typically 150,000 bytes, derived from a reference 100 Gbps link at 12 µs). The new CWND is calculated by adding a fraction of this constant to the current window:

CWND(new) = CWND(Init) + Base_BDP\Scaling Factor

Using a universal constant ensures Scale-Invariance in a mixed-speed fabric (e.g., 100G and 400G NICs).

If a 400G NIC were to use its own BDP (300,000 bytes) for the increase step, its window would grow four times faster than that of a 100G NIC. By using the shared Base_BDP (150,000 bytes), both NICs increase their throughput by the same number of bytes per second. This "normalized acceleration" prevents faster NICs from starving slower flows during the capacity-seeking phase.

As illustrated in Figure 6-8, consider a flow with an Initial CWND of 75,000 bytes, a Base_BDP of 150,000 bytes, and a Scaling Factor of 1024:

Step Size = 150,000 / 1024 ≈ 146.5 bytes

New CWND = 75,000 + 146.5 = 75,146.5 bytes

Note: Scaling factors are ideally set to powers of 2 (e.g., 512, 1024, 2048, 4096, 8192) to allow the hardware to use fast bit-shifting operations instead of expensive division.

Higher factors (e.g., 8192): Result in smaller, smoother increments (high stability).

Lower factors (e.g., 512): Result in larger increments (faster convergence to link rate).

Figure 6-8: Increasing CWND.

Tuesday, 6 January 2026

UET Congestion Management: CCC Base RTT

Calculating Base RTT

[Edit: January 7 2026, RTT role in CWND adjustment process]

As described in the previous section, the Bandwidth-Delay Product (BDP) is a baseline value used when setting the maximum size (MaxWnd) of the Congestion Window (CWND). The BDP is calculated by multiplying the lowest link speed among the source and destination nodes by the Base Round-Trip Time (Base_RTT).

In addition to its role in BDP calculation, Base_RTT plays a key role in the CWND adjustment process. During operation, the RTT measured for each packet is compared against the Base_RTT. If the measured RTT is significantly higher than the Base_RTT, the CWND is reduced. If the RTT is close to or lower than the Base_RTT, the CWND is allowed to increase.

This adjustment process is described in more detail in the upcoming sections.

The config_base_rtt parameter represents the RTT of the longest path between sender and receiver when no other packets are in flight. In other words, it reflects the minimum RTT under uncongested conditions. Figure 6-7 illustrates the individual delay components that together form the RTT.

Serialization Delay: The network shown in Figure 6-7 supports jumbo frames with an MTU of 9216 bytes. Serialization delay is measured in time per bit, so the frame size must first be converted from bytes to bits:

9216 bytes × 8 = 73,728 bits

Serialization delay is then calculated by dividing the frame size in bits by the link speed. For a 100 Gbps link:

73,728 bits / 100 Gbps = 0.737 µs

Note: In a cut-through switched network, which is standard in modern 100 Gbps and above data center fabrics, the switch does not wait for the full 9216-byte frame to arrive before forwarding it. Instead, it processes only the packet header (typically the first 64–128 bytes) to determine the destination MAC or IP address and immediately begins transmitting the packet on the egress port. While the tail of the packet is still arriving on the ingress port, the head is already leaving the switch.

This behavior creates a pipeline effect, where bits flow through the network similarly to water through a pipe. As a result, when calculating end-to-end latency from a first-bit-in to last-bit-out perspective, the serialization delay is effectively incurred only once—the time required to place the packet onto the first link.

Propagation delay: The time it takes for light to travel through the cabling infrastructure. In our example, the combined fiber-optic length between Rank 0 on Node 1A and GPU 7 on Node A2 is 50 meters. Light travels through fiber at approximately 5 ns per meter, resulting in a propagation delay of:

50 m × 5 ns/m = 250 ns = 0.250 µs

Switching Delay (Cut-Through): The time a packet spends inside a network switch while being processed before it is forwarded. This latency arises from internal operations such as examining the packet header, performing a Forwarding Information Base (FIB) lookup to determine the correct egress port, and updating internal buffers and queues.

In modern cut-through switches, much of this processing occurs while the packet is still being received, so the added delay per switch is very small. High-end 400G switches exhibit cut-through latencies on the order of 350–500 ns per switch. For a path traversing three switches, the total switching delay sums to approximately:

3 × 400 ns ≈ 1.2 µs

Thus, even with multiple hops, switching delay contributes only a modest portion to the total Base RTT in 100 Gbps and above data center fabrics.

Forward Error Correction(FEC) Delay: Forward Error Correction (FEC) ensures reliable, “lossless” data transfer in high-speed AI fabrics. It is required because high-speed optical links can experience bit errors due to signal distortion, fiber imperfections, or high-frequency signaling noise.

FEC operates using data blocks and symbols. The outgoing data is divided into fixed-size blocks, each consisting of data symbols. In 100G and 400G Ethernet FEC, one symbol = 10 bits. For example, a 514-symbol data block contains 514 × 10 = 5,140 bits of actual data.

To detect and correct errors, the switch or NIC ASIC computes parity symbols from the data block using Reed-Solomon (RS) math and appends them to the block. The combination of the original data and the parity symbols forms a codeword. For example, in RS(544, 514), the codeword has 544 symbols in total, of which 514 are data symbols and 30 are parity symbols. Each symbol is 10 bits, so the 30 parity symbols add 300 extra bits to the codeword.

At the receiver, the codeword is checked: the parity symbols are used to detect and correct any corrupted symbols in the original data block. Because RS-FEC operates on symbols rather than individual bits, if multiple bits within a single 10-bit symbol are corrupted, the entire symbol is corrected as a single unit.

The FEC latency (or accumulation delay) comes from the requirement to receive the entire codeword before error correction can begin. For a 400G RS(544, 514) codeword:

• 544 symbols × 10 bits/symbol = 5,440 bits total

• At 400 Gbps, this adds a fixed delay of ~150 ns per hop

This delay is a “fixed cost” of high-speed networking and must be included in the Base RTT calculation for AI fabrics. The sum of all delays gives the one-way delay, and the round-trip time (RTT) is obtained by multiplying this value by two. The config_base_rtt value in figure 6-7 is the RTT rounded to a safe, reasonable integer.

Figure 6-7: Calculating Base_RTT Value.

Saturday, 3 January 2026

UET Congestion Management: Congestion Control Context

Congestion Control Context

Updated 5.1.2026: Added CWND computation example into figure. Added CWND cmputaiton into text.
Updated 13.1.2026: Deprectade by: Ultra Ethernet: Congestion Control Context

Initializing Congestion Control Context (CCC)

When the PDS Manager receives an RMA operation request from the SES layer, it first checks whether a suitable Packet Delivery Context (PDC) already exists for the JobID, destination FEP, traffic class, and delivery mode carried in the request. If no matching PDC is found, the PDS Manager allocates a new one.

For the first PDC associated with a particular destination, a Congestion Control Context (CCC) is required to manage end-to-end congestion for that flow. The PDS Manager requests a CCC from the CCC Manager within the Congestion Management Sublayer (CMS). The CCC Manager creates the CCC, which initially enters the IDLE state, containing only the basic data structures without an active configuration. After creation, the CCC is bound to the PDC.

Next, the CCC is assigned a congestion window (CWND), which is computed based on CCC configuration parameters. The first step is to compute the Bandwidth-Delay Product (BDP), which is used to derive the upper bound for the initial congestion window. The CWND limits the total number of bytes in flight across all paths between the sender and the receiver.

The BDP is computed as:

BDP = min(sender_link_speed, receiver_link_speed) × config_base_rtt

The link speed must be expressed in bytes per second, not bits per second, because BDP is measured in bytes. The min() operator selects the smaller of the sender and receiver link speeds. In an AI fabric, these values are typically identical. The sender link speed, receiver link speed, and config_base_rtt are pre-assigned configuration parameters.

UET typically allows a maximum in-flight volume of 1.5 × BDP to provide throughput headroom while minimizing excessive queuing. A factor of 1.0 represents the minimum required to “fill the pipe” and would set the BDP directly as the maximum congestion window (MaxWnd). However, the UET specification applies a factor of 1.5 to allow controlled oversubscription and improved utilization.

Once the CWND is assigned and the CCC is bound to the PDC, the CCC transitions from the IDLE state to the ACTIVE state. In the ACTIVE state, the CCC holds all configuration information and is associated with the PDC, but data transport has not yet started.

When the CCC is fully configured and ready for operation, it transitions to the READY state. This transition signals that the CCC can enforce congestion control policies and monitor traffic. At this point, the PDC is allowed to begin sending data, and the CCC tracks and regulates the flow according to the configured congestion control algorithms.

The CCC serves as the central control structure for congestion management, hosting either sender-side (NSCC) or receiver-side (RCCC) algorithms. A CCC is unidirectional and is instantiated independently on both the sender and the receiver, where it is locally associated with the corresponding PDC. Once in the READY state, the CCC maintains the state required to regulate data flow, enabling NSCC and RCCC to enforce congestion windows, credits, and path usage to prevent network congestion and maintain efficient data transport.

Note: In this model, the PDS Manager acts as the control-plane authority responsible for context management and coordination, while the Packet Delivery Context (PDC) performs data-plane execution under the control of the Congestion Control Context (CCC). Once the CCC is operational and the PDC is authorized for data transport, RMA data transfers proceed directly over the PDC without further involvement from the PDS Manager.

Figure 6-6: Congestion Context: Initialization.