Saturday, August 1, 2026

Prinsengracht 263-265

The subtle geometry of the canal house holds steady, even as the world warps around it.
© 2026 Bryan R. Hinton
Provenance · Integrity Record
Hashes are of the byte-identical JPEGs converted from raw sensor data. Verify with sha512sum -c SHA512SUMS.
Fingerprint: CA42 47E8 9A5E FEAB 36DC 6A42 C547 9171 B69A 3CFB 887D B92C 3FB1 480A 2993 57A3
The .ots file proves that SHA512SUMS existed at or before the Bitcoin block timestamp below.
Anchor: Bitcoin block 954537
Timestamp: 2026-06-20 09:42 UTC
SHA512SUMS.ots

KL Auschwitz II-Birkenau · The Outer Perimeter

KL Auschwitz II-Birkenau. This is a primary record of my visit.
© 2026 Bryan R. Hinton

I took this photo while walking around the perimeter of Auschwitz II-Birkenau in Oświęcim, Poland. The barracks in the distance are to the left of the barracks where Edith, Margot, and Anne Frank stayed, specifically sector BIb of the women's camp.

After Birkenau, I traveled through Kraków and stopped to pick up some food and drinks for the long ride ahead, huddled up in the rail car.

Provenance · Integrity Record
Hash is of the byte-identical JPEG converted from raw sensor data. Verify with sha512sum -c SHA512SUMS.
Fingerprint: 2969 AEB8 0A42 E021 5E66 4D93 BD9B C85E 0213 8166 415C 9171 2074 3657 B386 AF51
The .ots file proves that SHA512SUMS existed at or before the Bitcoin block timestamp below.
Anchor: Bitcoin block 951065
Timestamp: 2026-05-26 02:49 UTC
SHA512SUMS.ots

Wednesday, July 29, 2026

Places of Memory · Konzentrationslager Auschwitz I · Between Block 10 and 11

Oświęcim, Poland

Auschwitz I and Auschwitz II-Birkenau are in Oświęcim, Poland which is about an hour from Kraków.

I took these photos at Auschwitz I.

These brick buildings comprise the majority of Auschwitz I. There are 28 of them. They are referred to as blocks. Many of them have underground areas.
© 2026 Bryan R. Hinton
Block 11 was a punishment block and the basement contains different types of punishment cells. I will not be going into any detail regarding the punishment methods and activities that occurred in these areas as it is sensitive material.
© 2026 Bryan R. Hinton

The discussion of the things that occurred within this block at Auschwitz I is sensitive to many people. But just as important, there is an equal or greater amount of dismissal, denial, and distortion of what really happened. Remembrance is important, and an acute awareness of what is going on now, in and across our world, is the only thing that has ever given us a chance of preventing it from happening again.

If you can't make it to Poland, I recommend visiting the United States Holocaust Memorial Museum in Washington, DC. The traveling exhibition Auschwitz. Not long ago. Not far away., produced by Musealia with the Auschwitz-Birkenau State Museum, is coming there. Here is the press release page.

The Wall of Death between Block 10 and Block 11 at Auschwitz I
© 2026 Bryan R. Hinton
Suitcase with the name "Anna Kraus" behind glass at Auschwitz I
© 2026 Bryan R. Hinton
Suitcase with the name "Sara" behind glass at Auschwitz I
© 2026 Bryan R. Hinton

Systematic renaming of Jewish individuals and families was written into German law before the war. A decree of 17 August 1938 forced Jews in Germany and annexed Austria to add "Sara" or "Israel" to their given names on every official document. The persecution continued through the war and its aftermath. These people were told they would begin a new life with work. Here is the luggage they brought with them on the train.

provenance · integrity record
Hashes are of the byte-identical JPEGs converted from raw sensor data. Verify with sha512sum -c SHA512SUMS
fingerprint: 2969 AEB8 0A42 E021 5E66 4D93 BD9B C85E 0213 8166 415C 9171 2074 3657 B386 AF51
The .ots file proves that SHA512SUMS existed at or before the Bitcoin block timestamp below.
Anchor: Bitcoin block 951183
Timestamp: 2026-05-26 23:40 UTC
SHA512SUMS.ots

Wednesday, July 8, 2026

Places of Memory · Kraków, Poland

Kraków, Poland
© 2026 Bryan R. Hinton
Provenance · Integrity Record
Hashes are of the byte-identical JPEGs converted from raw sensor data. Verify with sha512sum -c SHA512SUMS.
Fingerprint: 2969 AEB8 0A42 E021 5E66 4D93 BD9B C85E 0213 8166 415C 9171 2074 3657 B386 AF51
The .ots file proves that SHA512SUMS existed at or before the Bitcoin block timestamp below.
Anchor: Bitcoin block 951191
Timestamp: 2026-05-27 00:49 UTC
SHA512SUMS.ots

Wednesday, April 22, 2026

Die Ungebrochene Identität: Quantensichere Resistenz

Gedächtnis bedeutet, die Unveränderlichkeit der Wahrheit über die Zeit zu gewährleisten. In der physischen Welt nutzen wir Archive, um unsere Geschichten zu bewahren. In der digitalen Welt verwenden wir Kryptografie, um Identität, Urheberschaft und Vertrauen zu schützen.

Eine neue Bedrohung durch Quantencomputer fordert nun diese Grundlage heraus. In großem Maßstab wird sie in der Lage sein, die kryptografischen Aufzeichnungen zu löschen oder zu fälschen, die unser digitales Leben prägen.

Um die Integrität des kollektiven Gedächtnisses zu schützen und zu verhindern, dass zukünftige Angreifer Identitäten stehlen, habe ich frühere kryptografische Standards hinter mir gelassen und implementiere heute die höchste verfügbare Sicherheitsstufe: Post-Quanten-Technologie. Die doppelte Bedrohung: Shor und Grover

Quantencomputing stellt zwei unterschiedliche mathematische Bedrohungen für moderne Kryptografie dar. Um den Übergang zu Post-Quanten-Standards zu verstehen, ist es essenziell, beide zu kennen.

Shors Algorithmus: Der Public-Key-Zerstörer

Shors Algorithmus stellt die existenzielle Bedrohung dar. Er löst effizient die Probleme der ganzzahligen Faktorisierung und diskreten Logarithmen, die fast alle klassischen Public-Key-Kryptosysteme stützen, einschließlich RSA, Diffie-Hellman und elliptischen Kurven (ECC). Dies ist keine Schwächung, sondern ein vollständiger Bruch. Ein ausreichend leistungsfähiger Quantencomputer kann private Schlüssel aus öffentlichen Schlüsseln ableiten und untergräbt damit fundamentale Identitätssysteme.

Grover's Algorithmus: Der symmetrische Komprimierer

Grover's Algorithmus zielt auf symmetrische Kryptografie und Hash-Funktionen ab. Er bietet eine quadratische Beschleunigung für Brute-Force-Suchen und halbiert effektiv die Sicherheitsstärke eines Schlüssels. Daher ist AES-256 so entscheidend: Selbst nach Grovers Reduktion bietet es noch 128 Bit effektive Sicherheit, die praktisch unknackbar sind.

Die praktische Konsequenz: Jetzt speichern, später entschlüsseln

Die unmittelbarste Gefahr ist der SNDL-Angriff (Store Now, Decrypt Later). Verschlüsselter Datenverkehr, Identitätsnachweise, Zertifikate und Signaturen können heute abgefangen werden, während klassische Kryptografie noch gültig ist, und unbegrenzt gespeichert werden. Sobald Quantentechnologie ausgereift ist, können diese Archive nachträglich entschlüsselt oder gefälscht werden. Wenn unsere kryptografischen Grundlagen versagen, verlieren wir auch die Fähigkeit, unsere eigene digitale Geschichte zu dokumentieren.

Jenseits veralteter Standards: Warum ML-DSA-87

Jahrelang war elliptische Kurvenkryptografie, insbesondere P-384 (ECDSA), der Goldstandard in Hochsicherheitsumgebungen. Während P-384 etwa 192 Bit klassische Sicherheit bietet, hat es keinerlei Widerstand gegen Shors Algorithmus. Es wurde für eine klassische Welt entwickelt, und diese Welt geht zu Ende.

Daher habe ich ML-DSA-87 für Root-CA- und Signieroperationen implementiert. ML-DSA-87 ist die höchste Sicherheitsstufe moderner gitterbasierter Standards (Kategorie 5), rechnerisch äquivalent zu AES-256. Die Wahl dieser Stufe statt des verbreiteten ML-DSA-65 stellt sicher, dass die Identität meines Netzwerks mit dem heute größtmöglichen Sicherheitsspielraum aufgebaut ist.

Hardwarerealität: AArch64 und die PQC-Last

Post-Quanten-Kryptografie ist nicht länger theoretisch. Sie ist jetzt einsetzbar, sogar auf Routern und Mobilgeräten. Ich betreibe einen angepassten OpenSSL-3.5.0-Build auf einer AArch64 MediaTek Filogic 830/880-Plattform. Dieser SoC ist ungewöhnlich gut für Post-Quanten-Workloads geeignet.

Vektorskalierung mit NEON

ML-KEM und ML-DSA basieren stark auf Polynomarithmetik. ARM-NEON-Vektorbefehle ermöglichen die parallele Ausführung dieser Operationen und reduzieren so die TLS-Handshake-Latenz selbst bei großen PQ-Schlüsselmaterialien erheblich.

Speichereffizienz

Post-Quanten-Schlüssel sind groß. Ein öffentlicher ML-KEM-1024-Schlüssel umfasst 1568 Bytes, verglichen mit 49 Bytes für P-384. Der 64-Bit-Adressraum von AArch64 ermöglicht eine effiziente Verwaltung dieser Puffer und vermeidet Fragmentierungsprobleme älterer Architekturen.

Technische Verifikation: Post-Quanten-CLI-Prüfungen

Nach Installation des angepassten Toolchains auf dem AArch64-Zielsystem kann der Post-Quanten-Stack direkt verifiziert werden.

KEM-Verifikation

openssl list -kem-algorithms

Erwartete Ausgabe:

ml-kem-1024
secp384r1mlkem1024 (high-security hybrid)

Signaturverifikation

openssl list -signature-algorithms | grep -i ml

Erwartete Ausgabe:

ml-dsa-87 (256-bit security)

Das Vorhandensein dieser Algorithmen bestätigt, dass die Plattform sowohl Post-Quanten-Schlüsselaustausch (ML-KEM-1024) als auch quantenresistente Signaturen (ML-DSA-87) unterstützt.

Zusammenfassung: Mein AArch64-Post-Quanten-Stack

  • Bibliothek: OpenSSL 3.5.4 (angepasster AArch64-Build)
  • SoC: MediaTek Filogic 830 / 880
  • Architektur: ARMv8-A (AArch64)
  • Schlüsselaustausch: ML-KEM-1024 + Hybride
  • Identität & Signatur: ML-DSA-87
  • Sicherheitsstufe: Stufe 5 (quantenbereit)
  • Status: Produktionsreif

Durch den direkten Wechsel zu ML-KEM-1024 und ML-DSA-87 habe ich die veralteten Engpässe des letzten Jahrzehnts umgangen. Mein Netzwerk bereitet sich nicht mehr auf den Quantenübergang vor - es hat ihn bereits abgeschlossen. Der Rest der Industrie wird folgen.

Wednesday, January 12, 2022

Concurrency, Parallelism, and Barrier Synchronization - Multiprocess and Multithreaded Programming

On preemptive, timed-sliced UNIX or Linux operating systems such as Solaris, AIX, Linux, BSD, and OS X, program code from one process executes on the processor for a time slice or quantum. After this time has elapsed, program code from another process executes for a time quantum. Linux divides CPU time into epochs, and each process has a specified time quantum within an epoch. The execution quantum is so small that the interleaved execution of independent, schedulable entities – often performing unrelated tasks – gives the appearance of multiple software applications running in parallel.

When the currently executing process relinquishes the processor, either voluntarily or involuntarily, another process can execute its program code. This event is known as a context switch, which facilitates interleaved execution. Time-sliced, interleaved execution of program code within an address space is known as concurrency.

The Linux kernel is fully preemptive, which means that it can force a context switch for a higher priority process. When a context switch occurs, the state of a process is saved to its process control block, and another process resumes execution on the processor.

A UNIX process is considered heavyweight because it has its own address space, file descriptors, register state, and program counter. In Linux, this information is stored in the task_struct. However, when a process context switch occurs, this information must be saved, which is a computationally expensive operation.

Concurrency applies to both threads and processes. A thread is an independent sequence of execution within a UNIX process, and it is also considered a schedulable entity. Both threads and processes are scheduled for execution on a processor core, but thread context switching is lighter in weight than process context switching.

In UNIX, processes often have multiple threads of execution that share the process's memory space. When multiple threads of execution are running inside a process, they typically perform related tasks. The Linux user-space APIs for process and thread management abstract many details. However, the concurrency level can be adjusted to influence the time quantum so that the system throughput is affected by shorter and longer durations of schedulable entity execution time.

While threads are typically lighter weight than processes, there have been different implementations across UNIX and Linux operating systems over the years. The three models that typically define the implementations across preemptive, time-sliced, multi-user UNIX and Linux operating systems are defined as follows - 1:1, 1:N, and M:N where 1:1 refers to the mapping of one user-space thread to one kernel thread, 1:N refers to the mapping of multiple user-space threads to a single kernel thread. M:N refers to the mapping of N user-space threads to M kernel threads.

In the 1:1 model, one user-space thread is mapped to one kernel thread. This allows for true parallelism, as each thread can run on a separate processor core. However, creating and managing a large number of kernel threads can be expensive.

In the 1:N model, multiple user-space threads are mapped to a single kernel thread. This is more lightweight, as there are fewer kernel threads to create and manage. However, it does not allow for true parallelism, as only one thread can execute on a processor core at a time.

In the M:N model, N user-space threads are mapped to M kernel threads. This provides a balance between the 1:1 and 1:N models, as it allows for both true parallelism and lightweight thread creation and management. However, it can be complex to implement and can lead to issues with load balancing and resource allocation.

Parallelism on a time-sliced, preemptive operating system means the simultaneous execution of multiple schedulable entities over a time quantum. Both processes and threads can execute in parallel across multiple cores or processors. Concurrency and parallelism are at play on a multi-user system with preemptive time-slicing and multiple processor cores. Affinity scheduling refers to scheduling processes and threads across multiple cores so that their concurrent and parallel execution is close to optimal.

It's worth noting that affinity scheduling refers to the practice of assigning processes or threads to specific processors or cores to optimize their execution and minimize unnecessary context switching. This can improve overall system performance by reducing cache misses and increasing cache hits, among other benefits. In contrast, non-affinity scheduling allows processes and threads to be executed on any available processor or core, which can result in more frequent context switching and lower performance.

Software applications are often designed to solve computationally complex problems. If the algorithm to solve a computationally complex problem can be parallelized, then multiple threads or processes can all run at the same time across multiple cores. Each process or thread executes by itself and does not contend for resources with other threads or processes working on the other parts of the problem to be solved. When each thread or process reaches the point where it can no longer contribute any more work to the solution of the problem, it waits at the barrier if a barrier has been implemented in software. When all threads or processes reach the barrier, their work output is synchronized and often aggregated by the primary process. Complex test frameworks often implement the barrier synchronization problem when certain types of tests can be run in parallel. Most individual software applications running on preemptive, time-sliced, multi-user Linux and UNIX operating systems are not designed with heavy, parallel thread or parallel, multiprocess execution in mind.

Minimizing lock granularity increases concurrency, throughput, and execution efficiency when designing multithreaded and multiprocess software programs. Multithreaded and multiprocess programs that do not correctly utilize synchronization primitives often require countless hours of debugging. The use of semaphores, mutex locks, and other synchronization primitives should be minimized to the maximum extent possible in computer programs that share resources between multiple threads or processes. Proper program design allows schedulable entities to run parallel or concurrently with high throughput and minimum resource contention. This is optimal for solving computationally complex problems on preemptive, time-sliced, multi-user operating systems without requiring hard, real-time scheduling.

Wednesday, February 24, 2021

A hardware design for variable output frequency using an n-bit counter

The DE1-SoC from Terasic is an excellent board for hardware design and prototyping. The following VHDL process is from a hardware design created for the Terasic DE1-SoC FPGA. The ten switches and four buttons on the FPGA are used as an n-bit counter with an adjustable multiplier to increase the output frequency of one or more output pins at a 50% duty cycle.

As the switches are moved or the buttons are pressed, the seven-segment display is updated to reflect the numeric output frequency, and the output pin(s) are driven at the desired frequency. The onboard clock runs at 50MHz, and the signal on the output pins is set on the rising edge of the clock input signal (positive edge-triggered). At 50MHz, the output pins can be toggled at a maximum rate of 50 million cycles per second or 25 million rising edges of the clock per second. An LED attached to one of the output pins would blink 25 million times per second, not recognizable to the human eye. The persistence of vision, which is the time the human eye retains an image after it disappears from view, is approximately 1/16th of a second. Therefore, an LED blinking at 25 million times per second would appear as a continuous light to the human eye.

scaler <= compute_prescaler((to_integer(unsigned( SW )))*scaler_mlt);
gpiopulse_process : process(CLOCK_50, KEY(0))
begin
if (KEY(0) = '0') then -- async reset
count <= 0;
elsif rising_edge(CLOCK_50) then
if (count = scaler - 1) then
state <= not state;
count <= 0;
elsif (count = clk50divider) then -- auto reset
count <= 0;
else
count <= count + 1;
end if;
end if;
end process gpiopulse_process;
The scaler signal is calculated using the compute_prescaler function, which takes the value of a switch (SW) as an input, multiplies it with a multiplier (scaler_mlt), and then converts it to an integer using to_integer. This scaler signal is used to control the frequency of the pulse signal generated on the output pin.

The gpiopulse_process process is triggered by a rising edge of the CLOCK_50 signal and a push-button (KEY(0)) press. It includes an asynchronous reset when KEY(0) is pressed.

The count signal is incremented on each rising edge of the CLOCK_50 signal until it reaches the value of scaler - 1. When this happens, the state signal is inverted and count is reset to 0. If count reaches the value of clk50divider, it is also reset to 0.

Overall, this code generates a pulse signal with a frequency controlled by the value of a switch and a multiplier, which is generated on a specific output pin of the FPGA board. The pulse signal is toggled between two states at a frequency determined by the scaler signal.

It is important to note that concurrent statements within an architecture are executed concurrently, meaning that they are evaluated concurrently and in no particular order. However, the sequential statements within a process are executed sequentially, meaning that they are evaluated in order, one at a time. Processes themselves are executed concurrently with other processes, and each process has its own execution context.