Skip to main content

TECH VEDA

Linux kernel & Device drivers starts on 24th Oct 2026 enrollingCorporate on-site training - Submit proposal Pick your modulesSharpen your kernel skills: deep dives, drivers, Yocto, CVEs, careers — updated daily. Read the blog →Embedded Linux fast track starts 30th sept 2026 enrollingEmbedded Linux Mastery track starts 30th sept 2026 enrollingLinux systems engineering starts 30th sept 2026 enrolling
System Design

Application Processor Plus MCU: When a Two-Processor Design Is Worth It

When is a two-processor design worth it? A domain-map test, four options from hardware interlock to discrete MCU, and the numbers you must measure.

Application Processor Plus MCU: When a Two-Processor Design Is Worth It

Before adding a second processor, read your SoC’s power-domain map. On parts like AM62x and i.MX 93 the on-die real-time core sits in a domain that stays alive with the application processor powered down, which removes the two requirements most often used to justify a separate chip. A discrete MCU earns its place in a two-processor design when the domain map genuinely cannot meet the requirement, when field IO and isolation argue for it, when you need a common-cause-failure argument for certification, or when a certified MCU already exists from the last product. And before any of that: check whether the safe state can be held by hardware that has no firmware to fail.

The decision is usually framed as “Linux cannot do hard real time, so we need an MCU”. That framing skips two steps and gets the two-processor design question wrong in both directions — teams add a chip they did not need, and teams omit a hardware interlock they did. It is also hard to reverse, because it lands in the schematic, the BOM and the enclosure.

The context: three requirements that get confused

Deterministic control. A commutation loop, a switching pattern, a sensor sampled on a fixed period. The worst case matters, not the average.

Availability independent of Linux. Linux takes seconds to boot and reboots on update. If a heater, valve or brake must be held safe across that window, it cannot depend on Linux being up.

Low-power standby. Something stays awake for a button, an accelerometer or an RTC alarm while the rest of the system is powered down.

Only the first is about speed. The other two are about independence, and they are routinely collapsed into the first — which is how a two-processor design gets justified by a fast MCU solving an availability problem that a pull-down resistor would have solved better.

Run the domain-map test first

This step decides most of the argument, and it is skipped almost universally because the SoC block diagram does not show it. Open the datasheet at the power-domain and low-power-modes chapters and answer five questions about the on-die real-time core:

  1. Is it in an independently supplied power domain, or does it share rails with the application cores?
  2. Does it have an independent reset source, and is it on a separate reset hierarchy from the A-cluster warm reset?
  3. Can it run with the main domain powered off, and which wake sources remain live in that mode?
  4. Does its wake and run path avoid DDR, or does it need DRAM that the application processor will put into self-refresh?
  5. Can its code and hot data live in TCM or on-chip SRAM, and how much is there?

The answers are frequently better than engineers expect. TI documents an MCU Only mode on AM62x in which the MCU core stays alive running application code while the main domain is powered off and DDR sits in self-refresh, with MCU UART, MCU MCAN and MCU Timer available as wake sources; the MCU GPIO domain is documented as one that is never fully shut down. i.MX 93’s Cortex-M33 real-time domain works on the same principle. On a part like that, “it must survive a Linux reboot” and “it must sit in a standby state the application processor cannot reach” are no longer reasons to add a chip. On a part without an independent domain, they still are. The test tells you which world you are in, and nothing else about a two-processor design can be settled until you know.

Option 0 — Hardware default-safe, no processor

Before any processor argument: can the safe state be held by hardware that has no firmware? Pins Hi-Z on reset with an external pull to the safe level, a gate driver whose default is shutdown, UVLO, a latching comparator interlock, a window-watchdog supervisor IC with an enable line. Sub-dollar, no second build, and it has no software failure mode at all — which is why a safety assessor would rather see it than a processor.

Treat this as a precondition rather than an alternative. An MCU that holds an output safe is strictly worse than hardware that cannot fail to a dangerous state, so you do not get to claim the availability requirement until you have shown hardware cannot hold it.

Option 1 — One processor

Everything on the application processor, with a real-time scheduling policy and CPU isolation where timing matters. One toolchain, one image, one update — and worst-case timing bounded by the whole system rather than one loop.

The sub-option that gets missed sits here: programmable IO blocks — TI’s PRU-ICSS, PIO-style engines, high-resolution timer and FlexIO-class peripherals. These are architecturally different from a Cortex-M sharing an interconnect: cycle-deterministic, and on the PRU the firmware is loaded by remoteproc from the root filesystem, so there is no second update story and no second toolchain culture. The discriminator is the shape of the work, not its speed: fixed, tight, repetitive IO patterns go to programmable IO; stateful control with modes, protocol and decisions goes to an M-class core. A commutation pattern is usually the former, and teams escalate it to an MCU anyway.

Option 2 — Application processor with an on-die coprocessor

The industrial and automotive SoC families from NXP, TI and ST generally ship a real-time core on the same die — i.MX 8M Plus pairs four Cortex-A53 with a Cortex-M7 at 800 MHz, STM32MP1 pairs Cortex-A7 with a Cortex-M4 at 209 MHz. This is emphatically not true of the wider Linux SoC population: most Rockchip, Allwinner, Amlogic, Broadcom and Qualcomm application parts give you no usable Cortex-M, so this is a reason to shortlist those families early, not an assumption to carry into a BOM.

Two things here are commonly misread, and both matter more than the core’s clock speed.

The core is not always yours. On AM62x the Cortex-R5F is occupied: it runs TI’s Device Manager firmware, is started by the U-Boot R5 SPL early in boot, and ships inside tispl.bin. The application-facing core there is the Cortex-M4F. Read the vendor’s core-allocation documentation before counting cores on a block diagram.

Its determinism is not free. A Cortex-M fetching code or data from shared DDR inherits the application cluster’s memory traffic as jitter, and no amount of RTOS tuning fixes that. The working answer is to run the critical path entirely from TCM or on-chip SRAM — which caps firmware size and complexity, and is the real engineering cost of this option. That is what question 5 in the domain-map test is for.

What you get in exchange is supervision you would otherwise write yourself. Linux loads the core’s firmware, starts and stops it, and restarts it after a crash through the kernel’s remoteproc framework — all of which is yours to build across a board-level link. Know the limit before you count on it, though: remoteproc’s recovery is enabled by default and works, but nothing in the kernel detects a fault by itself. Recovery runs only when the platform driver reports a crash, which needs the SoC to raise an event, so a coprocessor stuck in a loop with no fault exception reads as running indefinitely. You get a restart primitive, not a liveness check, and closing that gap is your work either way.

Option 3 — Application processor plus a discrete MCU

This is what most people mean by a two-processor design: a separate MCU package with its own clock, supply and flash. Two variants are often conflated and should not be — on the same board over SPI or UART, and on a separate board over CAN or RS-485. The second is a different product, with a different cost, harness, EMC and field-replacement story.

You gain independence no domain map has to grant you, a part chosen on its merits, and a consolidation point for field IO. You pay in everything doubling — toolchain, build, signed image, update path, rollback, security advisories — and the supervision remoteproc gave you across a die now has to be designed into the schematic: an AP GPIO to the MCU’s reset line so Linux can restart it, a bidirectional heartbeat, and an external watchdog neither processor can defeat alone.

Two claims about discrete MCUs are repeated uncritically and deserve correcting. It does not boot in milliseconds in any sense that matters: reset to first instruction is microseconds, but crystal and PLL lock, a secure-boot signature check, sensor settling and safety start-up self-tests put a real product in the tens to low hundreds of milliseconds — and those self-tests are mandatory in exactly the certification case used to justify the MCU. In any case, for holding an output across a Linux reboot the MCU’s boot time is irrelevant because it never went down; the hardware default state from Option 0 is what holds it. And its standby current beats the application processor’s floor only at system level if the MCU can cut the AP’s rails, which means a PMIC enable line and a cold boot on wake — while board leakage and regulator quiescent current frequently dominate both figures anyway.

The numbers you have to produce

No article can give you the figures for your board, and any that offers them is guessing. The measurements that settle a two-processor design are specific, though, and a decision made without them is a preference:

  • Worst-case wake-to-response latency on the application processor, measured with cyclictest at your real thread priority while the system does what it will do in the field — display, network, storage, camera all active — for hours rather than minutes, because the tail is the point. The maximum decides Option 1; the average is irrelevant.
  • The same loop on the on-die core, run once from DDR and once entirely from TCM or SRAM. The gap between those two runs is the cost of question 5, and it is usually the surprise.
  • System standby current in each candidate mode, measured on your board with its regulators and leakage rather than from a datasheet summary — including the case where the MCU holds the AP’s rails off.
  • Time from power-on to first valid output on each candidate, which is a different requirement from surviving a reboot and is often the decisive one. Measure both and compare against your requirement: a core started from ROM or SPL will reach it far sooner than a Linux userspace can.
  • Cost, on both sides of the scale. Be careful here, because the two costs of a second processor sit on the same side: the MCU’s per-unit price and the permanent engineering cost of a second signed build, update path and advisory stream both count against Option 3, so volume never makes it cheaper. The real comparison is that combined figure against the engineering cost of making a single SoC meet the requirement — constraining firmware to fit SRAM, adding isolation on-board, or moving to a different part. There is no crossover volume to look for; there is a requirement, and the measurements above decide which side can meet it.

If you have to argue this rather than decide it — advising a client’s hardware team, or inheriting a frozen BOM — put one row per requirement and let the gaps show. An empty measurement cell is a more effective argument in a design review than any opinion:

RequirementWorst caseMeasured howMeets itAlternative costs
Valve closed during OTAMust hold indefinitelyInterrupt AP mid-updateOption 0—
Control loop deadline(your figure)cyclictest under load(blank until measured)(blank)
Standby current(your figure)Bench, per mode(blank)(blank)

One thing on that list cannot be deferred: peripheral and pin ownership between Linux and a coprocessor has to be divided before layout.

The decision: when a two-processor design earns its place

Start at Option 0 and make the design earn each step up. Run the domain-map test before comparing Options 2 and 3, because it moves two of the six criteria below. A discrete MCU is then justified when you can point at one of these and name the evidence:

  • Common-cause failure, or certified scope reduction. A core sharing a die shares its PLL, supply and errata, which is a common-cause problem; the more practical driver is keeping the safety function small so the Linux stack falls into a lower software safety class. Ask whether your vendor ships a safety package for the on-die core with an FMEDA and a safety manual — that single fact can settle this criterion either way. What IEC 61508, ISO 26262 or IEC 62304 require is a question for a functional-safety engineer, not an article.
  • Field IO topology, EMC and isolation. In industrial designs this is the most common real reason a discrete MCU exists, and it has nothing to do with determinism: you do not route 24 V field IO, relay drive, CAN and motor-phase sense into a fine-pitch BGA beside DDR. The MCU is an IO consolidation and isolation boundary, and it changes layer count, connector count and your EMC test outcome.
  • A standby state the application processor lacks, confirmed by question 3 — not by intuition.
  • Control that survives a Linux reboot, where neither Option 0 nor question 2 covers it. Note too that a coprocessor started by the bootloader and adopted through attach keeps running across a Linux restart on boot flows that support it.
  • Independent lifecycle and change-control scope. When updating Linux must not re-open a submission or a certificate, a physical boundary is what bounds that scope.
  • Brownfield: the MCU already exists. If the predecessor product has a certified MCU with firmware, test fixtures, a protocol and a supplier, the burden of proof inverts — Option 3 is the default and absorbing the function into Linux must earn the re-qualification cost.

If you cannot name which criterion applies, and cannot point at the clause in the domain map or the assessment plan behind it, you do not have the requirement yet — you have an intuition. Whichever way it goes, one hedge is worth taking: footprint the MCU, leave it unpopulated, and reserve and route the SPI or UART pins plus a spare interrupt and a reset line. Board area and a few hours of layout convert an irreversible decision into a deferrable one.

Worked example

A 24 V industrial valve controller on an AM62x: Modbus over RS-485, a touchscreen, remote firmware update, a 1 kHz position loop with a 200 µs deadline, and a spring-return valve that must go closed if anything fails. On the face of it a two-processor design; walk it through.

The five questions. Q1: the M4F sits in the MCU domain with its own supply. Q2: that domain has its own reset, and is not taken down by an A53 warm reboot — which is the question that decides criterion 4, so answer it explicitly rather than assuming. Q3: MCU Only mode keeps it alive with the main domain off. Q4: keep its working set off DDR. Q5: roughly 40 KB of code and hot data against the part’s on-chip SRAM, so the loop fits with room to spare — and that is a number to confirm on your own build, not to take from here.

The six criteria. Certification: the fail-closed state is answered by Option 0, because the spring return closes the valve with no firmware involved, which is a stronger claim than any processor can make. Field IO: this one has force — but what it argues for is an isolated 24 V stage, not a processor behind it. Standby and survives-a-reboot: Q2 and Q3 already answered both, so neither argues for a separate chip. Lifecycle and brownfield: greenfield, with no submission to re-open and no predecessor MCU. Result: one SoC, the M4F running the loop from SRAM, hardware default-safe on the valve, and an isolated 24 V IO stage — which the standing criterion actually argues for, rather than a processor behind it. Note what the framework really produced: the decisive move was the actuator specification, not the silicon. The measurements that would overturn it are a 200 µs worst case the M4F misses from SRAM, or an assessor rejecting a shared-die independence argument.

Consequences

Two consequences of a two-processor design belong in the decision itself, because they change which option you pick rather than merely how you build it.

A running coprocessor can cost the power saving that motivated it. A live core often blocks the application processor’s deeper idle states outright, so a second processor chosen for battery life can spend more system power than it saves. That is measurable before you commit — it is the third item in the measurement list.

Splitting late is worse than splitting early; merging back is worse still. Deciding at schematic stage costs a design review, discovering it after the first build costs a respin, and removing a discrete MCU once the control logic has settled on it rarely happens at all.

Everything else a split costs you belongs to the build rather than to this decision, and none of it should change which option you choose.

Key takeaways

  • Programmable IO beats an MCU for fixed tight IO patterns; an M-class core wins when the work is stateful.
  • Whichever way it goes, the decision is only as good as the measurement behind it — and the one measurement teams skip is the same loop run from DDR and from SRAM.
Was this worth your time?

Frequently asked questions

Does a PRU or programmable IO block count as a second processor for a safety argument?
Structurally it is weaker than a discrete MCU and stronger than a Linux thread. It shares the die, the supply and the errata with the application cores, so it does nothing for a common-cause-failure argument — but it does not share the scheduler, so it supports a scope-reduction argument well. If your assessor’s concern is shared silicon, a PRU does not answer it. If the concern is the size and complexity of the certified software, it can.

Should the MCU sit on the main board or on a separate one?
Follow the field IO. If the MCU exists mainly to keep 24 V wiring, relay drive and long cable runs away from the application processor, putting it on a separate module at the connector end serves that purpose better and shortens the runs that radiate. If it exists for timing or standby power, keep it on the main board — a field bus between the two adds latency and one more thing to certify.

Our assessor wants separate silicon, but the domain map says the on-die core is enough. Who wins?
The assessor, but find out which argument they are making first. If it is common-cause failure — shared PLL, supply and errata — the domain map cannot answer it and you need separate silicon. If it is about the Linux stack’s size and software safety class, an on-die core with the vendor’s safety package, FMEDA and safety manual may satisfy them. Ask which one it is before respinning the board.

Further reading

RB
Raghu Bharadwaj

Founder, TECH VEDA — 20+ years teaching the Linux kernel, device drivers and embedded systems.

Follow on LinkedIn

Get new posts by email

Kernel, embedded Linux and AI-era engineering — a few sharp reads a month. No spam.

We email occasionally and never share your address.