OSFEC.006: Direct Memory Access (DMA) — Moving Data Without Stalling the CPU

Comic-style technical illustration of DMA moving data from a peripheral to RAM while the CPU remains free.

Elementary Overview and Review

Direct Memory Access (DMA) is a hardware engine that moves data between memory and peripherals without making the CPU copy every byte itself. In OSFEC.002, we used tasks, priorities, interrupts, and timing to keep firmware responsive; in OSFEC.005, we built a safer firmware update path. DMA adds another production tool: let dedicated hardware move repetitive data while the processor keeps running control logic, protocol code, or an RTOS.

Embedded Systems Tutorials — a focused introduction to Direct Memory Access in embedded systems.

Why DMA Exists

Without DMA, firmware often polls a peripheral register or services an interrupt for each small piece of data. With DMA, the firmware configures a transfer, the peripheral generates requests, and the DMA controller becomes a temporary bus master that reads from one address and writes to another. STMicroelectronics’ DMA guide describes the same model: high-speed peripheral-to-memory, memory-to-peripheral, or memory-to-memory transfers happen with little CPU action, which frees processor time for other work.

Phil’s Lab — practical STM32 DMA with SPI and FreeRTOS, including transfer-complete handling.

Anatomy of a DMA Transfer

A DMA transfer is defined by a small set of parameters: source address, destination address, transfer count, element width, address-increment rules, request source, direction, and priority. If an ADC data register always lives at one memory-mapped register, the peripheral address stays fixed while the destination RAM address increments through a buffer. For a UART transmit operation, the opposite pattern is common: the RAM source increments while the UART data register stays fixed.

ControllersTech — register-level DMA setup with half-transfer and transfer-complete interrupts.

Circular Buffers Make Continuous Sampling Possible

For continuous streams such as ADC samples, audio, sensor data, or serial traffic, DMA can operate in circular mode. The controller fills a buffer, wraps to the beginning, and keeps going. Half-transfer and transfer-complete events let software process one region while DMA fills the other. This is the foundation of ping-pong and double-buffer designs: ownership must be clear so the CPU never edits a region that DMA is actively writing.

ControllersTech — STM32 ADC sampling from polling through interrupts and DMA, including continuous acquisition.

DMA, Interrupts, and RTOS Tasks Must Share Ownership Cleanly

DMA does not remove interrupts; it changes what they mean. Instead of interrupting the processor for every byte, firmware can receive one interrupt when a block is half full, complete, or in error. In an RTOS, the DMA interrupt handler should usually do the minimum work required to acknowledge the event and wake the appropriate task with a queue, semaphore, notification, or event flag. The task then processes the finished buffer outside interrupt context.

Eddie Amaya — practical STM32 DMA continuation showing how the transfer path is handled in firmware.

Cache Coherency Can Break Correct-Looking DMA Code

On processors with a data cache, especially Cortex-M7-class systems, the CPU and DMA engine may not immediately see the same bytes. The CPU can modify a cached buffer without writing it back to RAM, while DMA reads the older RAM copy; or DMA can update RAM while the CPU keeps reading stale cached data. Production firmware therefore needs an explicit policy: use non-cacheable DMA regions where appropriate, align buffers to cache-line boundaries, and perform the required cache clean or invalidate operations before ownership changes.

STMicroelectronics — Cortex-M7 data-cache behavior and coherency, directly relevant to DMA buffer ownership.

DMA Still Has Bandwidth and Latency Limits

DMA is not “free bandwidth.” The controller shares internal buses with the CPU and other masters, so arbitration, memory speed, peripheral timing, FIFO thresholds, burst size, and competing DMA channels all affect latency. ST’s AN4031 shows why firmware engineers must check whether the bus can sustain the required traffic. A simple first estimate is Required Bandwidth = samples per second × bytes per sample × channels; the real design then adds protocol overhead and safety margin.

EEVblog2 — STM32 DMA and ADC discussion focused on real data movement in an ARM microcontroller.

Minimal DMA Configuration Model

// Pseudocode: ADC peripheral -> RAM buffer

volatile uint16_t adc_samples[256];

void dma_start(void) {
    dma_disable(CHANNEL_ADC);

    dma_set_source(CHANNEL_ADC, &ADC_DATA_REGISTER);
    dma_set_destination(CHANNEL_ADC, adc_samples);
    dma_set_count(CHANNEL_ADC, 256);

    dma_set_width(CHANNEL_ADC, 16_bits);
    dma_source_increment(CHANNEL_ADC, false);
    dma_destination_increment(CHANNEL_ADC, true);

    dma_enable_half_transfer_irq(CHANNEL_ADC);
    dma_enable_transfer_complete_irq(CHANNEL_ADC);
    dma_enable_circular_mode(CHANNEL_ADC);

    dma_enable(CHANNEL_ADC);
    adc_enable_dma_requests();
}

DMA Buffer Ownership Pattern

DMA fills first half
        |
        v
Half-transfer interrupt
        |
        +----> CPU/task processes first half

DMA fills second half
        |
        v
Transfer-complete interrupt
        |
        +----> CPU/task processes second half

DMA wraps and repeats

Engineering Checklist

  1. Identify the exact peripheral request that triggers the DMA channel.
  2. Confirm source and destination addresses from the microcontroller reference manual.
  3. Set source/destination increment rules correctly.
  4. Match transfer width to the peripheral data width.
  5. Decide whether normal, circular, or double-buffer operation fits the workload.
  6. Keep DMA buffers alive for the entire transfer; never use a temporary stack buffer that disappears early.
  7. Define who owns each buffer region: DMA or CPU.
  8. Keep DMA interrupt handlers short and move processing into tasks where practical.
  9. Check cache maintenance requirements on cached cores.
  10. Calculate required bandwidth and allow margin for arbitration and competing traffic.
  11. Handle transfer error, FIFO error, underrun, overrun, and timeout cases.
  12. Use JTAG, SWD, OpenOCD, or GDB to inspect DMA registers, counters, buffers, and interrupt flags when the transfer stalls.

Exercises

  1. A 4-channel ADC samples each channel at 50 kS/s with 16-bit samples. Calculate the raw DMA bandwidth in bytes per second.
  2. For a peripheral-to-memory transfer, decide which address should increment and which should remain fixed.
  3. Draw a 512-sample circular buffer split into two 256-sample processing regions.
  4. Explain why a DMA buffer allocated inside a short-lived function can become unsafe after the function returns.
  5. Describe what a half-transfer interrupt allows the CPU to do before the full buffer is complete.
  6. Explain why cache invalidation may be required after DMA writes into RAM on a Cortex-M7 system.
  7. List three reasons a DMA channel can be configured correctly but still miss real-time deadlines.

Knowledge Check + Answers

  1. What is DMA? A hardware mechanism that moves data between memory and peripherals, or between memory regions, with little CPU involvement.
  2. Why is DMA useful? It reduces repetitive CPU copying and interrupt load while supporting high-throughput I/O.
  3. What does “fixed peripheral address” mean? The DMA repeatedly accesses the same hardware register while the memory address can advance through a buffer.
  4. What is circular DMA? A mode in which the DMA automatically wraps to the beginning of a buffer and continues transferring.
  5. What is a half-transfer interrupt? A notification that the first half of a programmed block is complete and can often be processed while DMA fills the second half.
  6. Does DMA eliminate interrupts? No. DMA commonly generates fewer, block-level interrupts such as half complete, complete, and error.
  7. Why can cache coherency matter? The CPU cache and RAM can temporarily contain different versions of the same buffer while DMA accesses RAM directly.
  8. Is DMA bandwidth unlimited? No. DMA still competes for buses, memory, and peripheral access and must satisfy real timing limits.
  9. What is the raw bandwidth for 4 × 50,000 samples/s × 2 bytes? 400,000 bytes per second, or about 400 kB/s before overhead.
  10. What is the most important ownership rule? The CPU and DMA should never modify the same buffer region at the same time unless the design explicitly guarantees that access is safe.

Remember

  • DMA moves data; the CPU still controls policy.
  • Configure addresses, widths, counts, triggers, and increment rules deliberately.
  • Circular DMA is ideal for continuous streams when buffer ownership is explicit.
  • Block-level interrupts reduce CPU overhead but still need correct synchronization.
  • Cached systems require a coherency strategy.
  • Always prove the required bandwidth and latency on real hardware.

Leave a comment