Showing posts with label Cisco Silicon ONE. Show all posts
Showing posts with label Cisco Silicon ONE. Show all posts

Monday, 2 March 2026

Packet trimming Deep Dive - Part IV

Receive Network Processing Unit (Rx NPU)

Figure 9-4 illustrates a simplified receive-side processing pipeline, starting from the moment a Packet Header Vector (PHV), constructed by the Rx IFG, is delivered to the Receive Network Processing Unit (Rx NPU).

When the PHV arrives at the Rx NPU, it is dispatched to one of the Run-to-Completion (RTC) cores in the Packet Processing Array (PPA). Each RTC core processes the packet within a single execution context, allowing parsing, classification, lookup, and queuing decisions to be resolved without intermediate handoffs between processing stages.

The first task of the RTC parser is to perform deep inspection of the packet headers. While the Rx IFG has already extracted basic Layer-2 and Layer-3 information, the RTC parser determines whether the packet is tunneled and whether the switch itself is the tunnel termination point. To demonstrate this behavior, consider a VXLAN-encapsulated packet. The outer Ethernet and IP headers are used to forward the packet through the underlay network. If the outer destination IP address matches one of the local switch IP addresses, the device identifies itself as the tunnel endpoint. The tunneling protocol is recognized by examining the UDP header, where destination port 4789 indicates VXLAN. After the tunneling mechanism is identified, the outer headers are logically removed, and processing continues using the inner Ethernet and IP headers. These inner headers then form the basis for forwarding decisions. In this example, tunneling illustrates a scenario in which the switch operates as a Virtual Tunnel Endpoint (VTEP) in a multitenant scale-out backend network.

In parallel with deep parsing, traffic classification takes place. The packet is assigned an Internal Traffic Class (ITC) by the pre-classification process in the Rx IFG pipeline. The ITC is used solely for internal prioritization within the Rx NPU, such as memory access arbitration and scheduling of processing resources inside the pipeline. It influences how the packet progresses through the internal stages of the NPU but does not determine where the packet is buffered for transmission. The ITC is carried as metadata within the Packet Header Vector (PHV), ensuring that every internal bus and memory controller treats the packet according to its pre-assigned urgency as it traverses the NPU.

The Rx NPU pipeline, in turn, matches the DSCP field against the configured QoS classification policy. The result of this policy evaluation determines the Virtual Output Queue (VOQ) into which the packet will be enqueued. As described in the VOQ chapter, VOQs are organized by Traffic Class and destination, ensuring that congestion affecting one egress interface does not introduce head-of-line blocking for traffic destined to another. This separation decouples ingress buffering from egress congestion and preserves fairness under load.

At the same time, the RTC core performs a forwarding lookup against the Forwarding Information Base (FIB). The lookup resolves the egress interface and any associated forwarding attributes. Once the egress interface is known, the previously selected VOQ is mapped to the corresponding Output Queue (OQ) on that interface. This mapping follows the Traffic Class–to–egress priority relationship described earlier, ensuring that packets maintain consistent priority semantics from ingress classification through egress scheduling.

After the VOQ-to-OQ mapping is established, the Traffic Manager initiates a credit request toward the Tx NPU scheduler. This interaction follows the credit-based flow control model described in the VOQ chapter. Conceptually, the request indicates that a packet of a given size, associated with a specific egress port and priority level, is ready for transmission. The scheduler evaluates whether the egress port’s microscopic FIFO has sufficient available buffer space and whether any higher-priority packets are waiting to be transmitted. If the conditions allow, credits equal to the packet length are granted.

Only after credits are granted is the packet permitted to move from Unified Shared Memory (USM) toward the egress pipeline. This strict separation between enqueueing and transmission prevents buffer overcommitment and enforces priority ordering during congestion. 

Once the Tx NPU scheduler grants the necessary credits, the packet is dequeued from the Unified Shared Memory and enters the Transmit Network Processing Unit (Tx NPU). Similar to the receive side, the transmit pipeline utilizes a Run-to-Completion (RTC) model within its own Packet Processing Array (PPA). This ensures that the final packet transformations are performed with the same deterministic, single-context efficiency as the initial ingress processing.

Upon entering the Tx NPU, the packet is dispatched to a Tx RTC core. The core's primary responsibility is header reconstruction and encapsulation. While the Rx NPU made the forwarding decision, the Tx NPU executes the "physical" rewrite. For a packet exiting a VTEP, this is where the RTC engine pushes the appropriate VXLAN, UDP, IP, and Ethernet headers onto the inner payload. Because this is a programmable RTC environment, the device can support complex, multi-label stacks, such as SRv6 or deep MPLS label impositions, without the "recirculation" penalties found in fixed-pipeline ASICs.

In addition to encapsulation, the Tx NPU performs a final round of Egress Policy Enforcement. This includes applying egress ACLs, updating packet counters for billing or monitoring, and inserting In-band Network Telemetry (INT) metadata if configured. This allows the switch to timestamp the packet at the precise moment of departure, providing nanosecond-accurate latency data.

The final stage of the Tx NPU involves the Output Queue (OQ) Scheduler. Even though the packet has already been "credited" for transmission, this local scheduler manages the final arbitration between different traffic classes sharing the same physical port. It ensures that a burst of low-priority bulk data does not jitter a high-priority data stream at the very last microsecond of the journey.

Finally, the fully formed packet is handed off to the MAC and PCS (Physical Coding Sublayer). Here, the digital data is serialized and mapped into PAM4 (Pulse Amplitude Modulation 4-level) symbols. These symbols are then modulated onto the physical medium, whether as electrical signals over a backplane or light pulses through an optical transceiver, completing the packet's journey through the Silicon One architecture.

In summary, the Rx NPU integrates tunnel awareness, forwarding lookup, QoS-based queuing, and credit-controlled admission into the egress pipeline within a single run-to-completion processing model. Internal Traffic Class governs how the packet is processed inside the NPU, while the QoS policy determines where the packet waits for transmission. This separation of responsibilities enables deterministic performance, scalable queuing, and strict priority enforcement across the switching fabric.


Figure 9-4: Rx NPU Pipeline.

Friday, 27 February 2026

Packet Trimming Deep Dive - Part III

Virtual Output Queue (VOQ)


The Silicon One VOQ Architecture

Instead of using dedicated deep interface buffers for packet queuing, Cisco Silicon One utilizes a Centralized Shared Memory architecture paired with a logical Virtual Output Queue (VOQ) mechanism. Because the VOQ concept is implemented within the Ingress (Rx) NPU entity, this queuing stage occurs after the initial ingress lookups but before the packet is switched across the internal fabric to the egress.

The VOQ model turns the traditional egress queuing model, where packets wait for serialization in a hardware buffer on the specific egress interface, upside down. While a VOQ is physically located on the ingress NPU, its ability to send traffic is controlled by the state of a small hardware Output Queue (OQ) on the egress interface.


Priority Mapping and Default State

As shown in Figure 9-3, a QoS policy can be created where a packet received on interface gi1/0/1 is assigned to Traffic Class 6 if the DSCP bits are set to EF (Expedited Forwarding). This configuration instantiates a VOQ specifically for that traffic class. In this hierarchy:

TC 7 (Control Plane/CS6): Mapped to OQ 1, the highest Strict Priority (Level 1).

TC 6 (DSCP-TRIMMED/EF): Mapped to OQ 2, the second-highest priority (Level 2).

By default, Silicon One enables VOQ 7 (Network Control) and VOQ 1 (Default/Best Effort). This ensures that critical control-plane traffic (DSCP 48/CS6) is guaranteed a high-priority path, to keep the network stable, while all unclassified data flows through the default VOQ. VOQs 5 – 2 are disabled by default.


Congestion Management and HoL Prevention

The VOQ mechanism is critical for handling "many-to-one" traffic patterns. For example, if four 800G ingress interfaces all send "elephant flows" to a single 800G egress interface, the egress OQ will become congested. Through a credit-based flow-control system, the egress port stops issuing credits to the specific ingress VOQs targeting it.

Because these queues are "virtualized" per output, this congestion does not cause Head-of-Line (HoL) blocking. Traffic destined for other, non-congested egress ports continues to receive credits and flow freely, even though they share the same physical ingress NPU.


Internal Isolation

The VOQ-OQ mechanism is entirely internal to the switch. Peer devices have no visibility into these internal queues; they only see the resulting serialized traffic stream and the DSCP/CoS markings in the packet headers.



Figure 9-3: Virtual Output Queue.

Wednesday, 25 February 2026

Packet Trimming Deep Dive - Part II

Receive Interface Group (Rx IFG)


Ingress Pre-Processing and Integrity

The Receive Interface Group (Rx IFG) is the ingress pre-processing stage that handles the incoming Ethernet bitstream before the packet enters the Packet Processing Array (PPA) of the Receive Network Processing Unit (Rx NPU) in the Cisco Silicon One architecture.

Processing begins at the Rx MAC. The Rx MAC reconstructs (“delimits”) the Ethernet frame from the Physical Coding Sublayer (PCS) bitstream and verifies frame integrity by computing a Frame Check Sequence (FCS) using the CRC-32 algorithm. If the computed FCS does not match the received FCS value, the frame is considered corrupted and is dropped immediately at ingress. If the CRC check succeeds, the frame is admitted for further processing. 

Shallow classification and Traffic Class mapping

After frame validation, the Rx IFG identifies the Ethernet MAC header and detects the presence of IEEE 802.1Q VLAN tags. The Rx IFG performs shallow classification to efficiently manage hardware resources before deeper protocol parsing and forwarding decisions are executed in the Rx NPU. When an IEEE 802.1Q VLAN tag is present, the Rx IFG extracts the Priority Code Point (PCP) bits from the VLAN tag and maps them to an Internal Traffic Class (ITC). Based on this mapping, the frame is placed into a specific port-based local hardware FIFO queue.

If the frame does not carry a VLAN tag, or if no usable CoS information is available, the frame is assigned to a default Internal Traffic Class and buffered in a standard port-based FIFO queue.

At this stage, no forwarding lookup is performed; the purpose is limited to frame validation and shallow classification.

Port-based FIFO queues and packet buffering

The small port-based FIFO queues serve two purposes.

First, they provide prioritization for writing packet data into shared SRAM. Prioritization is required because a large number of packets, originating from many ingress ports and representing thousands of concurrent flows, may arrive simultaneously at the Rx IFG. The FIFO queues regulate this contention and determine the order in which packets are admitted into shared memory based on their assigned Internal Traffic Class.

Second, in parallel with the memory write operation, the Rx IFG generates a compact packet metadata structure known as the Packet Header Vector (PHV).

Packet Header Vector (PHV) Creation

The Packet Header Vector summarizes essential information about the packet without requiring full packet inspection. It is created in the Rx IFG to offload basic packet characterization from the Rx NPU parser, thereby reducing processing overhead in the programmable pipeline and enabling higher sustained packet rates. The PHV includes metadata such as:


  • Ingress interface and timestamp, and Packet length
  • Pointer to the packet’s cell chain in memory
  • Internal Traffic Class (ITC), VLAN ID, and EtherType
  • Router MAC hit indication


The Rx IFG operates as a pattern-matching engine rather than performing linear, sequential parsing of the entire packet. For example, when the EtherType field indicates IPv4 (0x0800), the Rx IFG recognizes that the IP Protocol field, used to identify the transport protocol such as TCP or UDP, is located at a fixed offset within the IPv4 header. It can therefore extract that field directly without scanning the header byte by byte.

Similarly, when an IEEE 802.1Q VLAN tag is present, the Rx IFG accounts for the additional header field and adjusts the parsing offset accordingly to locate the correct EtherType position before proceeding with further field extraction.

The Router MAC field in the PHV is a hit indicator that signals whether the destination MAC address matches one of the switch’s configured router MAC addresses. A positive match indicates that the packet is addressed to the switch itself and allows the Rx NPU to immediately invoke Layer 3 processing logic, bypassing Layer 2 forwarding paths.

Handoff to the Rx NPU

Once the PHV is constructed, it is dispatched to one of the Run-to-Completion (RTC) cores in the Packet Processing Array (PPA) of the Rx NPU. Because the Rx IFG includes a Pre-Parser, it can perform flow-based hashing. It examines the extracted L2/L3 header fields, computes a hash value, and ensures that all packets belonging to the same flow are directed to the same PPA core.

This preserves packet ordering while distributing the overall processing load across the chip. In Silicon One, this mechanism is commonly referred to as Local Target ID (LTID) generation. The IFG assigns a target ID to each packet, instructing the hardware which processing “lane” the packet should follow



Figure 9-2: Receive Interface Group Processes.