<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Engineering Logs]]></title><description><![CDATA[A technical log on Advanced Technologies for next-gen constrained systems. Bridging R&D theory and industrial reality to design secure architectures, and break down complex engineering challenges.]]></description><link>https://youssefattia.hashnode.dev</link><generator>RSS for Node</generator><lastBuildDate>Mon, 05 Oct 2026 06:38:10 GMT</lastBuildDate><atom:link href="https://youssefattia.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[A Deep Dive into CUDA-Q and the Mølmer-Sørensen Gate]]></title><description><![CDATA[There’s a massive gap between implementing a quantum circuit on paper and executing a reliable gate on a physical device. While the textbook defines the Mølmer-Sørensen (MS) gate as a magical black bo]]></description><link>https://youssefattia.hashnode.dev/a-deep-dive-into-cuda-q-and-the-m-lmer-s-rensen-gate</link><guid isPermaLink="true">https://youssefattia.hashnode.dev/a-deep-dive-into-cuda-q-and-the-m-lmer-s-rensen-gate</guid><dc:creator><![CDATA[Youssef ATTIA]]></dc:creator><pubDate>Mon, 10 Aug 2026 22:58:33 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6973e898fd66ec50455f4e13/7f2afdae-c71d-4d57-b927-db3376f8d468.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>There’s a massive gap between implementing a quantum circuit on paper and executing a reliable gate on a physical device. While the textbook defines the Mølmer-Sørensen (MS) gate as a magical black box that generates entanglement, the reality is a brutal optimization problem.</p>
<p>In trapped-ion quantum computing, the MS gate relies on a precise balance between a coupling strength (χ) and an interaction time (t). You need χ⋅t≈π/8 to get the perfect Bell state. However, in the real world, your laser intensity drifts, your ion frequency shifts, and your pulse timings are never perfect.</p>
<p>In this article, I'm building a <strong>robust calibration pipeline</strong> using NVIDIA's CUDA-Q platform. I’ll walk you through how I leveraged GPU-accelerated quantum dynamics and classical optimization to find the perfect parameters for a high-fidelity MS gate.</p>
<h2>Project Overview :</h2>
<p>Here is exactly what you get out of the box:</p>
<ul>
<li><p><strong>Effective MS Hamiltonian</strong>: Simulates the core trapped-ion interaction using H=χ(Sx)^2 with Sx=σx0+σx1​.</p>
</li>
<li><p><strong>Classical Optimizer (L-BFGS-B)</strong>: Jointly tunes χ and t to reach the target geometric phase.</p>
</li>
<li><p><strong>Fidelity Proxy</strong>: A custom figure of merit based on computational-basis populations (P00​ and P11​) that hits exactly 1.0 at the ideal Bell state.</p>
</li>
<li><p><strong>High-Resolution Diagnostics</strong>: Generates time traces of fidelity, populations, and accumulated phase for deep visualization.</p>
</li>
<li><p><strong>GPU-Native Acceleration</strong>: Leverages the CUDA-Q dynamics backend (Runge-Kutta integrator) and <code>cupy</code> to run fast on NVIDIA GPUs.</p>
</li>
<li><p><strong>Cloud-Ready</strong>: Optimized to run seamlessly on <strong>Google Colab with a Tesla T4</strong> (no supercomputer required).</p>
</li>
</ul>
<h2>The Physics: The Effective MS Hamiltonian</h2>
<p>The magic of the MS gate lies in the interaction Hamiltonian. We are working in the "effective" picture where we treat the spin-spin interaction as:</p>
<p>H=χ(Sx​)^2</p>
<p>Where Sx​=σx0​+σx1​ (the Pauli-X operator applied to both qubits). This Hamiltonian drives the system from the ∣00⟩ state to the entangled Bell state [1/sqrt(2)].(∣00⟩+i∣11⟩).</p>
<p>The trick is that when χt=π/8, the populations of the computational basis states become perfect:</p>
<ul>
<li><p>P00​≈0.5</p>
</li>
<li><p>P11≈0.5</p>
</li>
<li><p>P01≈P10≈0</p>
</li>
</ul>
<p>This is the target of our optimization.</p>
<h2>The Code Architecture:</h2>
<p>When approaching this calibration problem, I had three non-negotiable design principles:</p>
<ol>
<li><p><strong>GPU Acceleration:</strong> Quantum dynamics simulations are computationally intensive. I used <code>cupy</code> and CUDA-Q's <code>RungeKuttaIntegrator</code> to run the time evolution on the GPU. This allows for fast, high-resolution trajectory analysis.</p>
</li>
<li><p><strong>Log-Space Optimization:</strong> The parameter space (χ and t) spans multiple orders of magnitude. I optimized in <em>log-space</em> (using <code>np.exp(x)</code>). This prevents the optimizer from getting stuck in local minima and allows it to scale effectively across different physical regimes.</p>
</li>
<li><p><strong>Soft Constraints:</strong> Instead of hard-coded constraints, I built a multi-objective cost function. This balances fidelity, phase accuracy, and pulse duration (we want the fastest gate possible to avoid decoherence).</p>
</li>
<li><p>Here is the core of the cost function logic:</p>
</li>
</ol>
<pre><code class="language-plaintext">def cost_function(x):
    chi = np.exp(x[0])
    t_final = np.exp(x[1]

    # Soft box constraints
    if not (1e3 &lt; chi &lt; 1e7) or not (5e-7 &lt; t_final &lt; 2e-5):
        return 10.0

    # The actual quantum simulation
    fid_final = simulate(chi, t_final)

    # Phase penalty (must hit exactly pi/8)
    phase_err = (chi * t_final - phi_target)**2

# Time penalty (we prefer a ~3 µs gate to minimize decoherence)
    time_pen = 0.02 * (t_final * 1e6 - 3.0)**2

    return (1.0 - fid_final) + 3.0 * phase_err + time_pen
</code></pre>
<h2>The Optimization Process :</h2>
<p>I used the <strong>L-BFGS-B</strong> optimizer from SciPy. This quasi-Newton method is excellent for smooth, differentiable (or at least well-behaved) objective functions like ours. It respects bounds and converges significantly faster than gradient-free methods.</p>
<h3>The Convergence Criteria</h3>
<ul>
<li><p><strong>Fidelity:</strong> I defined a "fidelity proxy" as <code>1 - |P00-0.5| - |P11-0.5|</code>. When this hits 1.0, we are at the ideal Bell state.</p>
</li>
<li><p><strong>Phase:</strong> The optimizer must push <code>chi * t</code> to exactly <code>pi/8</code> (0.3927 rad).</p>
</li>
</ul>
<h2>The Results: Validating the Perfect Gate</h2>
<p>After running the optimizer, the results speak for themselves. Here is the exact console output from the script, hitting the theoretical limits of double precision:</p>
<pre><code class="language-plaintext">CUDA-Q MS Bell-state effective calibration
chi_opt / 2π = 2.500e+04 Hz
t_opt = 2.500 µs
χ·t = 0.392699 rad (target π/8 = 0.392699)
fidelity = 1.000000
P00 = 0.5000
P11 = 0.5000
optimizer reached the target phase and Bell-state populations, nfev = 57
</code></pre>
<p>It takes only 57 function evaluations to reach perfection. This proves that my cost function creates a smooth, convex-like optimization landscape, and the RK integrator on the GPU provides stable gradients for the L-BFGS-B algorithm to follow.</p>
<h2>The Hardware Beneath the Code:</h2>
<p>Every design decision in this code is a direct translation of a physical constraint:</p>
<h3>The Zero-Time Guard</h3>
<p>The <code>max(float(t_final), 1e-9)</code> guard corresponds to the minimum pulse length the Arbitrary Waveform Generator (AWG) can output, the acousto-optic modulator (AOM) cannot switch the laser intensity on and off instantaneously, so an optimizer attempting to evaluate a pulse of zero duration isn't generating a mathematical error but a physically meaningless request, and the guard simply represents the finite rise-time of the optical chain.</p>
<h3>The Log-Space Optimization</h3>
<p>The Rabi frequency and pulse length appear as theoretical parameters, but in the lab the laser intensity drifts by a few percent every hour due to thermal expansion of the optical table, a linear search space would incorrectly treat a drift from 10 kHz to 11 kHz identically to a drift from 1 MHz to 1.001 MHz, whereas physically relative drifts matter equally across all orders of magnitude, so optimizing in log-space forces the optimizer to treat all scales with equal importance, the exact logic behind tuning the RF power knobs in the control room.</p>
<h3>The Time Penalty</h3>
<p>The 3 µs bias in the cost function encodes the sweet spot for a typical Ytterbium-171 chain, staying too long fights against the 10-50 ms coherence times and more critically allows anomalous heating from trap electrodes to excite the motional mode, while going much below 1 µs breaks the Lamb-Dicke regime, exciting higher motional states and introducing off-resonant carrier transitions, so the <code>0.02 * (t_final * 1e6 - 3.0)**2</code> penalty embeds the physical trade-off between adiabaticity and motional heating directly into the cost landscape rather than leaving it as an arbitrary regularization weight.</p>
<h2>Conclusion: Bridging to the Solder</h2>
<h3>The Noisy Reality of Measurement:</h3>
<p>The cost function in the lab is a random variable constructed from finite shot statistics, where measuring P00​ and P11​ requires fluorescence state discrimination with variance scaling as 1/sqrt(N)​; each shot deposits energy into the motional mode and necessitates ground-state cooling before the next cycle, making every evaluation an expensive operation measured in hundreds of milliseconds rather than microseconds, so the optimizer cannot assume smooth deterministic gradients and must move to Gaussian-process surrogate models with acquisition functions that explicitly balance exploration against the trap's heating rate.</p>
<h3>The Drift That Never Stops</h3>
<p>A calibration converging in 2 minutes looks impressive on a simulator, but in the laboratory those 2 minutes represent a 0.1 °C temperature drift in the magnetic shield that shifts the qubit frequency by hundreds of Hertz through the second-order Zeeman effect, while the Raman beam pointing drifts by tens of microradians as the optical table settles, so the optimizer cannot terminate and declare success, it must run continuously as a background tracking loop, initializing each cycle at the previous optimum and following the slow drift of the laboratory environment like a PID controller chasing a moving setpoint.</p>
<p><a href="https://github.com/attiayoussef-TA/cudaq-ms-calibration">Github repository</a></p>
]]></content:encoded></item><item><title><![CDATA[The Day Encryption Dies: Engineering 'PANTHEON', A Hybrid Quantum-Classical Defense]]></title><description><![CDATA[By Youssef ATTIA
In the context of modern cryptographic defense, "Harvest Now, Decrypt Later" is not a theoretical threat, it is an architectural constraint. Nation-state actors are actively capturing encrypted traffic, banking on the eventual maturi...]]></description><link>https://youssefattia.hashnode.dev/the-day-encryption-dies-engineering-pantheon-a-hybrid-quantum-classical-defense</link><guid isPermaLink="true">https://youssefattia.hashnode.dev/the-day-encryption-dies-engineering-pantheon-a-hybrid-quantum-classical-defense</guid><dc:creator><![CDATA[Youssef ATTIA]]></dc:creator><pubDate>Sun, 25 Jan 2026 01:22:06 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1769299515612/7312b3cf-6eb9-4558-a562-07c5c0ca198d.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>By Youssef ATTIA</strong></p>
<p>In the context of modern cryptographic defense, <strong>"Harvest Now, Decrypt Later"</strong> is not a theoretical threat, it is an architectural constraint. Nation-state actors are actively capturing encrypted traffic, banking on the eventual maturity of Fault-Tolerant Quantum Computers (FTQC) to break RSA and ECC primitives.</p>
<p>This vulnerability is catastrophic for <strong>IoT and Edge networks</strong>. unlike transient web sessions, edge devices, autonomous vehicles, smart grid sensors, and industrial controllers operate on decade-long lifecycles. Millions of devices deployed today are hard-coded with encryption standards that will be broken before the hardware is decommissioned. We are effectively building smart cities with expiration dates.</p>
<p>To explore the defense mechanisms against this threat, I engineered <strong>PANTHEON</strong>: a comprehensive simulation of a hybrid cryptographic network. It integrates <strong>Quantum Key Distribution (QKD)</strong> for physical layer security with <strong>Post-Quantum Cryptography (PQC)</strong> for authentication, bridging the gap between quantum physics and classical network engineering.</p>
<hr />
<h2 id="heading-1-system-architecture-the-why-of-hybrid">1. System Architecture: The "Why" of Hybrid</h2>
<p>A common critique in crypto engineering is: <em>"Why not just use Post-Quantum Cryptography (PQC) and call it a day?"</em></p>
<p>The answer lies in <strong>Forward Secrecy</strong>. PQC is mathematical; it relies on the hardness of lattice-based problems (like Shortest Vector Problem). If a mathematical shortcut is discovered in 20 years, today's PQC traffic is compromised.</p>
<p><strong>QKD is physical.</strong> It relies on the laws of quantum mechanics. If an eavesdropper (Eve) attempts to intercept the key distribution, the <strong>No-Cloning Theorem</strong> dictates that she <em>must</em> introduce errors into the system. This allows Alice and Bob to detect the intrusion <em>before</em> the key is ever used.</p>
<p>PANTHEON implements a hybrid stack:</p>
<ol>
<li><p><strong>Physical Layer (Simulated):</strong> <a target="_blank" href="https://qiskit.org/">Qiskit</a> (Stabilizer method) simulates the qubit transmission and collapse.</p>
</li>
<li><p><strong>Link Layer:</strong> A custom implementation of the <a target="_blank" href="https://cascade-python.readthedocs.io/en/latest/protocol.html">Cascade Protocol</a> handles error correction (reconciliation).</p>
</li>
<li><p><strong>Network Layer:</strong> <a target="_blank" href="https://networkx.org/">NetworkX</a> manages topology and Dijkstra routing.</p>
</li>
<li><p><strong>Application Layer:</strong> <a target="_blank" href="https://openquantumsafe.org/">LibOQS</a> provides Dilithium2 signatures to authenticate the classical discussion, while AES-GCM handles the payload.</p>
</li>
</ol>
<hr />
<h2 id="heading-2-the-physical-layer-simulating-entanglement-and-collapse">2. The Physical Layer: Simulating Entanglement and Collapse</h2>
<p>The foundation of the simulation is the <strong>BB84 protocol</strong>. Alice prepares qubits in either the Z (computational) or X (Hadamard) basis. Bob measures them randomly.</p>
<p>I utilized <code>Qiskit</code> to model this behavior. The challenge was performance and topology constraints: simulating statevectors for thousands of qubits is computationally expensive, and the backend coupling map enforces a strict limit (<strong>maximum 14 qubits</strong>).</p>
<h3 id="heading-implementation-the-circuit">Implementation: The Circuit</h3>
<p>To respect these <strong>limits</strong> while generating <strong>long keys</strong> (4096+ bits), the system processes the stream in "chunks" (batches). Here is the core circuit logic for a single batch:</p>
<p>Python</p>
<pre><code class="lang-plaintext">def bb84_run(num_bits: int):
    # Randomly choose bits and bases
    bits = [random.randint(0, 1) for _ in range(num_bits)]
    sender_bases = [random.choice(['Z', 'X']) for _ in range(num_bits)]

    qc = QuantumCircuit(num_bits, num_bits)
    for i in range(num_bits):
        if bits[i] == 1: qc.x(i)           # Prepare |1&gt;
        if sender_bases[i] == 'X': qc.h(i) # Rotate to Hadamard basis

        # ... Channel simulation (Eve/Noise injection) ...

        if receiver_bases[i] == 'X': qc.h(i) # Rotate back for measurement

    qc.measure(range(num_bits), range(num_bits))

    # The Optimization: Stabilizer method
    backend = AerSimulator(method='stabilizer') 
    compiled_qc = transpile(qc, backend=backend, optimization_level=0)
</code></pre>
<h3 id="heading-the-gotcha-why-methodstabilizer-matters">The "Gotcha": Why <code>method='stabilizer'</code> matters</h3>
<p>A standard statevector simulation scales exponentially with qubit count (2^n). If you try to simulate 4096 bits this way, your RAM will vanish. However, BB84 only utilizes Clifford gates (Hadamard, X, Z). The <strong>Clifford group</strong> can be simulated on classical hardware in polynomial time (O(n^2)) using the stabilizer formalism (Gottesman-Knill theorem). By explicitly setting <code>method='stabilizer'</code>, I optimized PANTHEON to simulate thousands of bits per second on a consumer laptop.</p>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769302919710/b9c62798-7ee6-4fd5-b501-398e674ccaf0.png" alt class="image--center mx-auto" /></p>
<p><strong>Figure 1:</strong> Qiskit circuit for a single batch of qubits between Bob and Eve. You can see the random preparation gates (X, H), resets, and final measurements.</p>
<hr />
<h2 id="heading-3-the-reconciliation-layer-implementing-cascade">3. The Reconciliation Layer: Implementing Cascade</h2>
<p>Physical quantum channels are noisy. Even without Eve, photon detectors have dark counts. This results in the "Sifted Keys" of Alice and Bob having a <strong>Quantum Bit Error Rate (QBER)</strong>. They match <em>mostly</em>, but not <em>exactly</em>.</p>
<p>To fix this without revealing the key, I implemented the <strong>Cascade Protocol</strong>. This is an interactive parity-check protocol that performs a binary search to locate errors.</p>
<h3 id="heading-implementation-recursive-binary-search">Implementation: Recursive Binary Search</h3>
<p>Python</p>
<pre><code class="lang-plaintext">def cascade_reconciliation(kA, kB, max_passes=4):
    # ... Block splitting logic ...
    for start in range(offset, n, block_size):
        end = min(start + block_size, n)
        # Parity check: Does the sum of bits (mod 2) match?
        if sum(A[start:end]) % 2 != sum(B[start:end]) % 2:
            # Error detected! Begin Binary Search to locate index.
            l, r = start, end
            while r - l &gt; 1:
                mid = (l + r) // 2
                # Check left half parity recursively
                if sum(A[l:mid]) % 2 != sum(B[l:mid]) % 2:
                    r = mid
                else:
                    l = mid
            # Flip the erroneous bit on Bob's side
            B[l] = 1 - B[l]
</code></pre>
<h3 id="heading-the-gotcha-interactive-latency">The "Gotcha": Interactive Latency</h3>
<p>While Cascade is efficient at correcting errors near the Shannon limit, it is "chatty." It requires multiple round-trip times (RTTs) between Alice and Bob to exchange parity bits. In a simulation, this is instant. In a real-world scenario over fiber optics, this RTT latency is the primary bottleneck for key generation rates. This implementation highlights why modern production systems are moving toward <em>LDPC (Low-Density Parity-Check)</em> codes, which are non-interactive.</p>
<hr />
<h2 id="heading-4-privacy-amplification-the-mathematical-guarantee">4. Privacy Amplification: The Mathematical Guarantee</h2>
<p>Even after correcting errors, we must assume that every parity bit revealed during Cascade was intercepted by Eve. Furthermore, Eve may have measured some qubits directly.</p>
<p>To handle this, we perform <strong>Privacy Amplification</strong>. We compress the key by hashing it, reducing its length by an amount proportional to Eve's potential knowledge.</p>
<h3 id="heading-implementation-entropy-bounding">Implementation: Entropy Bounding</h3>
<p>Python</p>
<pre><code class="lang-plaintext">def privacy_amplification(key_bits, qber_est, leaked_bits, sec_param=40):
    n = len(key_bits)
    # Calculate Binary Entropy of the error rate
    h2_qber = -qber_est * math.log2(qber_est) - (1-qber_est) * math.log2(1-qber_est)

    # Secure Length Bound:
    # m = n * (1 - h2(QBER)) - (bits revealed in Cascade) - (safety parameter)
    bound = math.floor(n * (1.0 - h2_qber))
    m = bound - leaked_bits - sec_param

    if m &lt;= 0: return None # Security cannot be guaranteed

    # Compress using SHA-256
    return hashlib.sha256(bits_to_bytes(key_bits)).digest()[:m]
</code></pre>
<hr />
<h2 id="heading-5-network-integration-amp-experimental-results">5. Network Integration &amp; Experimental Results</h2>
<p>Finally, the physics simulation is wrapped in a <code>NetworkX</code> graph. The system routes packets using Dijkstra’s algorithm, and each hop uses a distinct QKD-derived key for AES-GCM encryption, signed with Dilithium2.</p>
<h3 id="heading-51-configuration">5.1. Configuration</h3>
<p>You must configure the environment variables before running the simulation to dynamically define the quantum channel properties and security constraints.</p>
<h3 id="heading-52-simulation-metrics">5.2. Simulation Metrics</h3>
<p>The simulations were performed on a graph consisting of five nodes (Alice, Bob, Charlie, David, and Eve).</p>
<ul>
<li><p><strong>Throughput:</strong> For each run, a total of <strong>4096 bits</strong> is generated and processed in successive batches (292 circuits of 14 qubits each).</p>
</li>
<li><p><strong>Noise Models:</strong> Background channel noise is modeled with an error rate of 0.01.</p>
</li>
<li><p><strong>Eve's Impact:</strong> Interference from Eve is configurable. When set to an eavesdropping fraction of 10%, the <strong>QBER</strong> spikes visibly.</p>
</li>
<li><p><strong>Abort Threshold:</strong> Any key is automatically discarded if the QBER exceeds the theoretical safety threshold of <strong>11%</strong>.</p>
</li>
<li><p><strong>Output:</strong> Key lengths before and after Privacy Amplification are recorded, with the minimum acceptable key size determined by the <code>PA_OUT_BITS_MIN</code> environment variable (defaulting to 64 bits).</p>
</li>
</ul>
<hr />
<h2 id="heading-6-final-thoughts">6. Final Thoughts</h2>
<p>PANTHEON demonstrates that building quantum-safe networks is not just a physics problem, it is a systems engineering problem. While the QKD hardware (lasers and detectors) is complex, the software stack governing it reconciliation, privacy amplification, and classical authentication requires rigorous engineering standards.</p>
<p>By combining <code>Qiskit</code> for the quantum layer and <code>liboqs</code> for the classical authentication layer, we can simulate and validate the architectures that will secure the next decade of the internet.</p>
<hr />
<p><a target="_blank" href="https://github.com/attiayoussef-TA/Qryptum-Prototype">GitHub Repository</a></p>
]]></content:encoded></item><item><title><![CDATA[Architecting a Bare-Metal Separation Kernel: Enforcing ARINC 653 on Cortex-M4]]></title><description><![CDATA[By Youssef ATTIA
1.0 The Safety Case: Why Isolation Matters
In the world of avionics, "blue screens of death" are not an option. If the In-Flight Entertainment (IFE) system crashes, it cannot be allowed to take down the Landing Gear controller.
To so...]]></description><link>https://youssefattia.hashnode.dev/architectural-design-study-bare-metal-separation-kernel-arinc653</link><guid isPermaLink="true">https://youssefattia.hashnode.dev/architectural-design-study-bare-metal-separation-kernel-arinc653</guid><dc:creator><![CDATA[Youssef ATTIA]]></dc:creator><pubDate>Sat, 24 Jan 2026 02:34:06 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1769266123348/6f3fab63-b5f5-4a29-88df-55a907ee5587.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>By Youssef ATTIA</strong></p>
<h2 id="heading-10-the-safety-case-why-isolation-matters"><strong>1.0 The Safety Case: Why Isolation Matters</strong></h2>
<p>In the world of avionics, "blue screens of death" are not an option. If the In-Flight Entertainment (IFE) system crashes, it cannot be allowed to take down the Landing Gear controller.</p>
<p>To solve this, modern aircraft use <strong>IMA (Integrated Modular Avionics)</strong> architectures governed by standards like <a target="_blank" href="https://en.wikipedia.org/wiki/ARINC_653">ARINC 653</a>. The core concept is <strong>Partitioning</strong>: strict isolation in Time and Space.</p>
<p>For my latest engineering study, I decided to move away from theoretical OS concepts and <strong>conduct an architectural design study</strong> on a <strong>Zero-Trust Separation Kernel</strong> from scratch, targeting the specific hardware constraints of the <a target="_blank" href="https://www.st.com/en/microcontrollers-microprocessors/stm32f4-series.html">ARM Cortex-M4</a> (STM32F4).</p>
<p>Here is the architecture of the <strong>Aero-Partition-Hypervisor</strong>.</p>
<h2 id="heading-20-the-concept-the-dynamic-iron-curtain">2.0 The Concept: The "Dynamic Iron Curtain"</h2>
<p>Most simple RTOS schedulers just save and restore the CPU registers (Context Switching). My kernel goes a step further.</p>
<p>It implements a <strong>Dynamic MPU Reconfiguration</strong>. Every time the scheduler switches tasks, it doesn't just switch the Stack Pointer; it physically reprograms the <strong>Memory Protection Unit (MPU)</strong> to lock down the entire memory map, leaving only the active partition's memory accessible.</p>
<p>If Task A tries to peek at Task B's memory? <strong>Hardware Fault. Immediate Reset.</strong></p>
<h2 id="heading-30-the-architecture">3.0 The Architecture</h2>
<p><img src="https://cdn.hashnode.com/res/hashnode/image/upload/v1769218342358/c16a9074-653b-42a3-a2c0-3f464dabef17.png" alt class="image--center mx-auto" /></p>
<h2 id="heading-40-the-brain-deterministic-scheduling">4.0 The Brain: Deterministic Scheduling</h2>
<p>The heartbeat of the system is the <code>SysTick</code> timer, configured in <code>src/systick.s</code> to fire exactly every 1ms. This enforces the <strong>"Time Partitioning"</strong> requirement of ARINC 653.</p>
<p>But the real magic happens in the <strong>PendSV Handler</strong>. This is where we perform the context switch.</p>
<p>In a standard OS, this is 20 lines of assembly. In a Separation Kernel, we have to reconfigure the hardware trust model mid-flight. Here is the critical section from my <code>src/context_switch.s</code>:</p>
<p>Code snippet</p>
<pre><code class="lang-plaintext">/* --- RECONFIGURE MPU --- */
ldr r4, =MPU_RNR
mov r5, #4              /* Select Region 4 (User Slot) */
str r5, [r4]

ldr r5, [r2, #8]        /* Load Base Addr from TCB */
ldr r4, =MPU_RBAR
str r5, [r4]

ldr r5, [r2, #12]       /* Load Attributes from TCB */
ldr r4, =MPU_RASR
str r5, [r4]

dsb
isb                     /* Apply MPU changes immediately */
</code></pre>
<p><strong>Why this matters:</strong> We are loading <code>MPU_RBAR</code> (Base Address) and <code>MPU_RASR</code> (Attributes) directly from the <strong>Task Control Block (TCB)</strong>. This means every task carries its own <strong>"Access Visa."</strong> The moment <code>isb</code> (Instruction Sync Barrier) executes, the hardware wall is enforced.</p>
<h2 id="heading-50-the-engineering-challenge-faking-the-stack-frame">5.0 The Engineering Challenge: "Faking" the Stack Frame</h2>
<p>The hardest part of this project wasn't the MPU; it was the Context Switch initialization.</p>
<p>To start a task on Cortex-M4, you can't just jump to a function pointer. You have to trick the CPU into thinking it was <em>already</em> running that task and was interrupted. This requires manually fabricating a <a target="_blank" href="https://developer.arm.com/documentation">hardware stack frame</a> in RAM that matches exactly what the processor expects to pop during an exception return.</p>
<p>I had to reverse-engineer the exception entry sequence to build the loader in <code>src/task_init.s</code>:</p>
<p>Code snippet</p>
<pre><code class="lang-plaintext">/* Faking the Stack Frame */
mov r1, #0x01000000      /* xPSR: Set T-bit (Thumb Mode) */
str r1, [r0, #-4]!
ldr r1, =task_landing_gear
str r1, [r0, #-4]!       /* PC: The Task Entry Point */
mov r1, #0
str r1, [r0, #-4]!       /* LR: Link Register */
</code></pre>
<p><strong>The "Gotcha":</strong> If you forget to set the <strong>T-bit</strong> (Thumb Mode) in the <code>xPSR</code> (Program Status Register), the CPU attempts to switch to ARM mode. Since Cortex-M4 only supports Thumb instructions, this triggers an immediate <code>UsageFault</code>. Since the Cortex-M4 does not support ARM-state instructions, the <strong>hardware logic</strong> dictates that this manual stack fabrication must be perfect to allow the first context switch.</p>
<h2 id="heading-60-the-shield-build-time-safety-assertions">6.0 The Shield: Build-Time Safety Assertions</h2>
<p>Safety isn't just about runtime checks; it's about build-time guarantees. I modified the <code>linker.ld</code> script to include <strong>Safety Assertions</strong>.</p>
<p>This configuration enforces a <strong>static contract</strong>: the build system is designed to actively reject any binary where the Kernel segment exceeds its allocated flash or RAM partition, preventing memory overlaps before the code ever touches the chip.</p>
<p>Code snippet</p>
<pre><code class="lang-plaintext">/* linker.ld - Safety Assertions */
/* 1. Check if Kernel RAM overflowed */
ASSERT(SIZEOF(.data) &lt;= LENGTH(RAM_KERNEL), "Error: Kernel Data too large for RAM_KERNEL")

/* 2. Check if Code fits in Flash */
ASSERT(SIZEOF(.text.kernel) &lt;= LENGTH(FLASH_KERNEL), "Error: Kernel Code too large")
</code></pre>
<p>Additionally, the <code>Makefile</code> runs a static analysis pass using <code>-fstack-usage</code> to generate <code>.su</code> files, allowing us to audit stack depth before ever flashing the chip.</p>
<h2 id="heading-70-the-interface-safe-system-calls">7.0 The Interface: Safe System Calls</h2>
<p>User partitions run in <strong>Unprivileged Mode</strong>. They cannot disable interrupts or touch hardware directly.</p>
<p>To yield control, they must use the <code>SVC</code> (Supervisor Call) instruction. However, we do <strong>not</strong> switch tasks immediately inside the SVC Handler. Instead, we set the <strong>PendSV</strong> flag.</p>
<p>Code snippet</p>
<pre><code class="lang-plaintext">svc_yield:
    /* Trigger a PendSV to context switch asynchronously */
    ldr r0, =ICSR
    ldr r1, =PENDSVSET
    str r1, [r0]
    bx lr
</code></pre>
<p><strong>Why do this?</strong> By deferring the context switch to the <code>PendSV</code> exception (which has the lowest priority), we ensure that if a critical hardware interrupt (like an accelerometer data burst) arrives at the exact same moment, the hardware handles the interrupt <strong>first</strong>, and only performs the context switch afterwards. This guarantees <strong>minimal interrupt latency</strong>.</p>
<h2 id="heading-80-the-gateway-decoding-the-exception-frame">8.0 The Gateway: Decoding the Exception Frame</h2>
<p>The most dangerous moment in any OS is the transition from User Mode to Kernel Mode. How does the Kernel know <em>who</em> called it and <em>what</em> they want?</p>
<p>In <code>src/syscalls.s</code>, I implemented a robust <strong>SVC Handler</strong> that manually unwinds the stack frame.</p>
<p>When an exception occurs, the Cortex-M4 pushes the context onto <em>either</em> the <strong>Main Stack (MSP)</strong> or the <strong>Process Stack (PSP)</strong>. To know which one, we must inspect <strong>Bit 2</strong> of the Link Register (<code>LR</code>).</p>
<p>Code snippet</p>
<pre><code class="lang-plaintext">SVC_Handler:
    /* 1. Determine which Stack Pointer was used */
    tst lr, #4                  /* Check Bit 2 of EXC_RETURN */
    ite eq
    mrseq r0, msp               /* Bit 2 = 0: Came from Kernel (MSP) */
    mrsne r0, psp               /* Bit 2 = 1: Came from User (PSP) */

    /* 2. Locate the SVC Instruction */
    /* r0 now points to the Stack Frame. 
       The Saved PC is at offset 24 (6 words * 4 bytes). */
    ldr r1, [r0, #24]           /* Load the Return Address (PC) */
    ldrb r0, [r1, #-2]          /* Read byte BEFORE the PC */

    /* 3. Execute System Call */
    cmp r0, #0
    beq svc_yield               /* If SVC #0 -&gt; Yield */
</code></pre>
<p><strong>The Robustness Guarantee:</strong> Most tutorials hardcode <code>mrs r0, psp</code>. My kernel dynamically checks <code>lr</code> to ensure that even if a Kernel function triggers a syscall (which uses MSP), the system won't crash reading the wrong memory address. This makes the handler <strong>reentrant and safe</strong>.</p>
<h2 id="heading-90-conclusion">9.0 Conclusion</h2>
<p>By enforcing isolation at the hardware level (MPU), the build level (Linker), and the exception level (Stack Unwinding), we move from a "Trust" model to a "Zero-Trust" model.</p>
<p>While this kernel is architected for the STM32 microcontroller, the <strong>principles of Time and Space Partitioning</strong> it demonstrates are identical to those protecting millions of passengers on modern airliners.</p>
<hr />
<p><a target="_blank" href="https://github.com/attiayoussef-TA/Aero-partition-hypervisor"><strong>GitHub Repository</strong></a></p>
]]></content:encoded></item></channel></rss>