Elementary Overview and Review
A microcontroller normally starts from a fixed reset location, but a production product often needs a small bootloader to decide which application image is safe to run. The previous engineer lesson, OSFEC.004, showed how the linker script, startup code, Flash, RAM, and vector table create a runnable image. This lesson adds the update layer: how firmware is received, checked, selected, activated, tested, rolled back, and recovered without turning an ordinary update into a bricked device.
A Bootloader Is a Decision Layer Between Reset and the Application
At reset, the processor must eventually reach a valid application reset handler, but a bootloader can pause first to inspect image metadata, update flags, hardware state, and recovery requests. A robust design keeps the bootloader small, deterministic, and easier to verify than the full application. It may initialize only the hardware required for update and validation, then relocate the vector table, load the application’s initial stack pointer, and branch to the selected image. If the system uses an RTOS, the bootloader usually hands control over before the normal scheduler starts. Production bootloaders such as MCUboot formalize this idea with defined image formats, slots, validation rules, and upgrade state.
Transport and Update Policy Are Different Problems
The path that carries a firmware image is the transport; the logic that decides whether the image may run is the update policy. A device might receive bytes through UART, USB DFU, CAN, Ethernet, Wi-Fi or Bluetooth, yet the bootloader should still enforce the same checks before activation. During development, JTAG, SWD, OpenOCD, and GDB may bypass the normal update path entirely, which is why engineers must distinguish “I can program the chip” from “the deployed product can update itself safely.”
Every Candidate Image Needs Identity and Integrity Data
A bootloader should not treat a random block of Flash as valid firmware. A production image normally carries an image header or manifest containing fields such as magic value, image size, hardware compatibility, semantic version, build identifier, flags, and an integrity value. A CRC or checksum can detect accidental corruption, while a cryptographic hash such as SHA-256 gives a strong digest of the image contents. The bootloader compares calculated values with expected metadata before it changes boot state. This is the same evidence-first principle used when validating a golden image: verify the bytes before trusting them.
Integrity Is Not the Same as Authenticity
A CRC can prove that bytes changed, but it cannot prove who created the image. A secure update system therefore adds digital signatures: the release process hashes the firmware and signs the digest with a protected private key, while the bootloader verifies the signature with a trusted public key. MCUboot supports signed images and tools such as imgtool for producing the required header and trailer. This creates a chain of trust in which “valid bytes” is not enough—the device also asks whether the image was authorized. The private signing key belongs in controlled release infrastructure, not inside application source code, a developer laptop image, or the field device.
A/B Slots Let the Old Image Survive the Update
One of the safest common layouts keeps two application regions: an active slot A and an inactive slot B. The running firmware downloads the candidate into the inactive slot instead of erasing itself. After integrity and authenticity checks pass, the bootloader can mark the new slot as pending and boot it on the next reset. If the candidate reaches a defined healthy state, the application confirms it; if it crashes, resets repeatedly, or never confirms, the bootloader can roll back to the previous image. Dual-bank Flash on some microcontrollers can support this pattern directly, while other devices implement the same idea with partitions in one Flash array or external nonvolatile memory.
Rollback Must Be Designed Together With Watchdogs and Health Checks
A bootloader cannot know that a new image is healthy merely because it reached main(). Production firmware should define a confirmation point such as “scheduler running, critical peripherals initialized, persistent storage mounted, and communications alive.” Until then, the image remains provisional. A watchdog timer helps detect hangs or reset loops, and persistent boot counters can stop a device from retrying a broken image forever. Engineers must also decide which failures should trigger rollback, which should enter a service/recovery mode, and which data migrations are safe to reverse. A/B firmware is powerful only when boot state, application health, and recovery behavior are treated as one state machine.
Versioning Prevents the Wrong Image From Winning
A safe updater compares more than “new file available.” Firmware metadata should identify the product, board revision, image type, version, and any required bootloader or data-schema version. The bootloader may reject an image built for different hardware even if its signature is valid. Security-sensitive products may also enforce anti-rollback, refusing versions older than a stored security counter so a known-vulnerable release cannot be reinstalled. Operational rollback and security anti-rollback can conflict, so engineers must define the policy deliberately: keep enough recovery capability to survive a bad deployment without allowing an attacker to force the device back to obsolete trusted-but-vulnerable code.
OTA Changes the Delivery Path, Not the Safety Requirements
Over-the-air (OTA) updates add fleet delivery, network interruption, partial downloads, credentials, rollout groups, and remote observability, but the device-side safety rules remain familiar. Download into a non-active region, verify the full image, activate atomically, reboot into a provisional state, confirm health, and keep a recovery path. An OTA client may use network connectivity to fetch the image while the bootloader remains network-agnostic and validates only local Flash. This separation reduces complexity in the most privileged code. It also means a cloud deployment system should never be the only line of defense against a malformed, corrupted, or unauthorized firmware image.
Production Firmware Updates Need a Release Pipeline, Not Just Bootloader Code
The bootloader is only one part of the system. A production release should preserve the exact source revision, toolchain version, linker map, ELF, binary, image metadata, signature, signing policy, compatibility matrix, and test evidence. A host-side utility written in Python can package or transfer images, but it should not silently invent version numbers or bypass verification. Continuous integration should build deterministic artifacts, automated tests should exercise successful upgrade, interrupted transfer, corrupt image, bad signature, power loss, failed first boot, rollback, and manual recovery, and the final artifact identity should match what the fleet receives. This turns firmware updating from an ad-hoc maintenance action into a repeatable engineering process.
Boot Decision Pseudocode
on_reset():
load_boot_state()
if recovery_requested():
enter_recovery_mode()
candidate = select_pending_or_active_image()
if not header_is_sane(candidate):
candidate = known_good_image()
if not integrity_check(candidate):
candidate = known_good_image()
if secure_boot_enabled() and not signature_valid(candidate):
candidate = known_good_image()
if candidate_is_pending():
increment_boot_attempt_counter(candidate)
if boot_attempts_exceeded(candidate):
candidate = previous_confirmed_image()
jump_to_application(candidate)
MCUboot-Style Signing Example
# Example workflow; exact options depend on MCUboot/Zephyr configuration.
imgtool keygen -k production-signing-key.pem -t ecdsa-p256
imgtool sign \
-k production-signing-key.pem \
--version 1.5.0 \
--header-size 0x200 \
--slot-size 0x80000 \
application.bin \
application.signed.bin
Useful Firmware Image Fields
- Magic value: Quickly identifies the expected image format.
- Image length: Prevents reads beyond the valid image region.
- Product / board ID: Stops a valid image from booting on incompatible hardware.
- Version: Supports upgrade ordering, rollback policy, and fleet inventory.
- Build ID or source commit: Connects a field image to engineering evidence.
- Hash or CRC: Detects corruption.
- Signature: Proves authorization when secure boot is enabled.
- Flags: Can mark pending, confirmed, encrypted, compressed, or test images.
Flash Capacity Equation
For a simple internal-Flash A/B design, a useful first estimate is Total Flash Required ≈ Bootloader + 2 × Maximum Application Image + Metadata/Settings + Safety Margin. For example, a 128 KiB bootloader plus two 768 KiB application slots already consumes 1,664 KiB before settings, crash logs, manufacturing data, or spare sectors are counted. That is why the memory budget from OSFEC.004 must be revisited when update architecture changes: safe field updates may require nearly twice the application storage even though runtime RAM usage barely changes.
Engineering Checklist
- Define the reset-to-application boot sequence and keep the bootloader as small as practical.
- Partition Flash deliberately for bootloader, active image, candidate image, metadata, persistent settings, and recovery needs.
- Separate transport code from image-validation and activation policy.
- Validate image size, hardware compatibility, version, and memory boundaries before activation.
- Use CRC/checksum for accidental corruption detection and cryptographic signatures when authenticity matters.
- Protect private signing keys outside field devices and ordinary source repositories.
- Use pending/confirmed boot state instead of declaring success immediately after programming.
- Connect rollback policy to watchdog resets, boot counters, health checks, and persistent state.
- Test power loss during erase, program, verification, activation, and first boot.
- Test invalid header, truncated image, corrupted image, wrong product ID, bad signature, and downgrade attempts.
- Preserve a documented recovery path using an appropriate DFU, SWD/JTAG, SPI-programmer, or service interface.
- Archive the exact release artifacts and verify that the deployed binary matches the approved build.
Exercises
- Draw a Flash map containing a 128 KiB bootloader, two 640 KiB application slots, 32 KiB settings region, and 64 KiB recovery/log region.
- Explain why a CRC can detect corruption but cannot prove that firmware came from an authorized release team.
- Design a three-state image lifecycle using downloaded → pending → confirmed.
- Describe what the bootloader should do after three consecutive watchdog resets from a pending image.
- Explain why a board-revision field belongs in firmware metadata.
- List five moments during an update where unexpected power loss must be tested.
- Compare wired USB DFU with OTA delivery while keeping the same validation policy.
- Explain how anti-rollback can improve security while making emergency recovery more complicated.
Knowledge Check + Answers
- What is the bootloader’s core job? Select and start an application image according to defined validation, update, and recovery rules.
- What is the difference between transport and policy? Transport moves firmware bytes; policy decides whether those bytes may become executable firmware.
- Why use two image slots? The old confirmed image can remain intact while a new candidate is downloaded and tested.
- What does a CRC prove? It can detect accidental data corruption; by itself it does not authenticate the publisher.
- What does a digital signature add? It allows the device to verify that an authorized private key signed the firmware image.
- What is a pending image? A candidate selected for trial boot but not yet permanently accepted.
- When should an image be confirmed? After the application has reached a defined healthy operating state, not merely after reset.
- What is anti-rollback? A policy that prevents installation or boot of firmware below an allowed security/version floor.
- Why are boot counters useful? They let the bootloader detect repeated failed starts and choose rollback or recovery instead of looping forever.
- Why archive the ELF, map file, binary, metadata, and signature together? They let engineers reproduce, debug, verify, and prove the exact firmware release that was deployed.
Elementary Conclusion
Think of the bootloader as a careful gatekeeper standing between reset and the main program. It checks that the firmware belongs to this device, that the bytes arrived correctly, and—when secure boot is used—that the image was signed by someone the device trusts. A/B slots let the old working firmware stay nearby while the new version proves itself. Rollback, watchdog/reset logic, version checks, and a real recovery path turn that gatekeeper into a safety system. Once those pieces are engineered together, firmware updating stops being “erase the old code and hope” and becomes a controlled transaction that can survive corrupted downloads, interrupted power, bad releases, and field failures.

Leave a comment