Skip to main content

TECH VEDA

Linux kernel & Device drivers starts on 24th Oct 2026 enrollingCorporate on-site training - Submit proposal Pick your modulesSharpen your kernel skills: deep dives, drivers, Yocto, CVEs, careers — updated daily. Read the blog →Embedded Linux fast track starts 30th sept 2026 enrollingEmbedded Linux Mastery track starts 30th sept 2026 enrollingLinux systems engineering starts 30th sept 2026 enrolling
Deep Dives

From start_kernel() to Userspace: A Line-by-Line Walkthrough

Walk the path from start_kernel to userspace on arm64: rest_init, kernel_init, the initcall levels, and the fixed order the kernel searches for init.

From start_kernel() to Userspace: A Line-by-Line Walkthrough

On arm64 the kernel’s own startup is a short, fixed chain: primary_entry in head.S calls start_kernel(), which ends in rest_init(), which creates PID 1 running kernel_init(), which executes the first userspace program. Nearly every failure in this region lands in one of three places: before console_init(), where nothing is printed at all; inside do_basic_setup(), where one initcall never returns; or at the final kernel_execve(), where the kernel prints No working init found. The console log tells you which of the three you are in, and the order in which the kernel searches for an init binary is fixed and worth knowing exactly.

Most engineers can name the pieces of the boot path and still not say, under pressure, what the kernel was doing when their board went quiet. The region from start_kernel to userspace is where a bring-up either finishes or stalls, and its three failure modes look identical from outside: a console that stops. This walkthrough follows the real call chain in Linux 6.12 on arm64 and says which stage each console line proves.

Three places a board stops between the bootloader and the shell prompt

Scope first. The chain below was read in init/main.c, init/do_mounts.c and arch/arm64/kernel/head.S from Linux 6.12. The arm64 entry sequence is architecture-specific; everything from start_kernel() onwards is generic and reads the same on arm, riscv and x86. The function names and the init search order are unchanged in the 6.1 and 6.6 longterm series, so this applies to a vendor BSP on either.

A board can stop in three places, told apart by what has already been printed:

  • Nothing after the bootloader’s last line. The kernel has not reached console_init(), which sits late in start_kernel(). Everything before it is silent unless an early console is active.
  • The banner and a long log, then silence. The kernel is inside do_basic_setup() running driver initcalls, and one of them has not returned.
  • A full log ending in Freeing unused kernel memory, then a panic. The kernel finished its own startup and failed to execute a userspace program.

These are not interchangeable. The usual mistake is to read the first as a bad kernel image and the third as a failed root mount; both are wrong often enough to cost a day.

What calls start_kernel on arm64, and with what state

The arm64 image header branches to primary_entry, which records whether the MMU was already on, saves the bootloader’s registers, builds an identity map and sets the CPU up:

SYM_CODE_START(primary_entry)
	bl	record_mmu_state
	bl	preserve_boot_args
	...
	bl	init_kernel_el			// w0=cpu_boot_mode
	mov	x20, x0
	bl	__cpu_setup			// initialise processor
	b	__primary_switch
SYM_CODE_END(primary_entry)

Two details matter later. preserve_boot_args copies x0 into x21, and x0 at kernel entry is the physical address of the device tree blob. init_kernel_el decides whether the kernel runs at EL1 or EL2, which is why a kernel entered at the wrong exception level fails here, long before any driver runs.

__primary_switch then enables the MMU, maps and relocates the kernel, and jumps to the last assembly before C:

SYM_FUNC_START_LOCAL(__primary_switched)
	adr_l	x4, init_task
	init_cpu_task x4, x5, x6
	...
	str_l	x21, __fdt_pointer, x5		// Save FDT pointer
	...
	bl	finalise_el2			// Prefer VHE if possible
	ldp	x29, x30, [sp], #16
	bl	start_kernel
	ASM_BUG()
SYM_FUNC_END(__primary_switched)

So on entry to start_kernel() the MMU is on, the stack belongs to init_task — the statically allocated task that becomes PID 0 — the exception vector base is set, and the device tree pointer is in a global. The ASM_BUG() after the call is not decoration: start_kernel() is __noreturn, and if it ever returned, that is where the board would stop.

Inside start_kernel: four ordering constraints that matter

The body of start_kernel() is a flat list of roughly eighty calls. Reading it as a list is not useful; reading it as four constraints is.

	local_irq_disable();
	early_boot_irqs_disabled = true;
	boot_cpu_init();
	page_address_init();
	pr_notice("%s", linux_banner);
	setup_arch(&command_line);
	...
	pr_notice("Kernel command line: %s\n", saved_command_line);
	parse_early_param();
	...
	trap_init();
	mm_core_init();
	...
	sched_init();
	rcu_init();
	init_IRQ();
	tick_init();
	time_init();
	...
	early_boot_irqs_disabled = false;
	local_irq_enable();
	kmem_cache_init_late();
	console_init();
	...
	rest_init();

Interrupts are off across most of it. local_irq_disable() is near the top and local_irq_enable() only comes after time_init(). Anything that sleeps or waits for an interrupt in that window hangs, which is why early code polls and why a bad timer or interrupt-controller driver stops the board with no output.

There is no allocator until mm_core_init(). Before it, allocations go through memblock_alloc(). Several calls are placed where they are purely because they need large allocations early; setup_log_buf(0) and vfs_caches_init_early() both sit above mm_core_init() with a comment saying so.

console_init() is late, and that explains the first failure mode. Every line printed before it, including the banner and the command line, is buffered and appears only once a console registers. A kernel that dies in setup_arch() or mm_core_init() has the explanation sitting in the log buffer with nothing to read it out.

The command line is parsed in two passes. parse_early_param() runs the early_param() handlers, then parse_args() handles the rest, and anything still unrecognised goes to unknown_bootoption(), which forwards it to userspace. That produces a notice worth reading rather than ignoring:

	pr_notice("Unknown kernel command line parameters \"%s\", will be passed to user space.\n",
		&unknown_options[1]);

A parameter you expected the kernel to act on, appearing in that line, means the kernel did not consume it. The usual cause is a config symbol that is not enabled, so its __setup handler was never linked in.

rest_init: why PID 1 exists before the scheduler has finished

The last call in start_kernel() is rest_init(), where the three familiar early PIDs appear. It is short enough to read whole:

static noinline void __ref __noreturn rest_init(void)
{
	struct task_struct *tsk;
	int pid;

	rcu_scheduler_starting();
	/*
	 * We need to spawn init first so that it obtains pid 1, however
	 * the init task will end up wanting to create kthreads, which, if
	 * we schedule it before we create kthreadd, will OOPS.
	 */
	pid = user_mode_thread(kernel_init, NULL, CLONE_FS);
	...
	pid = kernel_thread(kthreadd, NULL, NULL, CLONE_FS | CLONE_FILES);
	...
	system_state = SYSTEM_SCHEDULING;

	complete(&kthreadd_done);

	/*
	 * The boot idle thread must execute schedule()
	 * at least once to get things moving:
	 */
	schedule_preempt_disabled();
	/* Call into cpu_idle with preempt disabled */
	cpu_startup_entry(CPUHP_ONLINE);
}

Three facts fall out, and each answers a question engineers ask about this stage.

PID 1 is created with user_mode_thread(), not kernel_thread(). It begins life as a kernel thread running kernel_init() and becomes a userspace process only later, when kernel_execve() replaces its image. The same task is both, which is why a failure to execute init is a problem inside PID 1 rather than a failure to create it.

The ordering is deliberate, and the comment says why. kernel_init is spawned before kthreadd so it wins PID 1, but kernel_init()‘s first statement is wait_for_completion(&kthreadd_done). PID 1 is created first and then immediately blocks until PID 2 exists, because what it does next may need kernel threads. That is the kind of decision that looks accidental until you read the comment beside it.

The boot CPU does not disappear. After creating both tasks, rest_init() calls cpu_startup_entry(CPUHP_ONLINE) and never returns: the context that ran start_kernel() becomes the idle task, PID 0.

Everything kernel_init_freeable must finish before init is looked for

kernel_init() calls kernel_init_freeable() first, and that is where system bring-up happens: the GFP mask is opened so blocking allocations are allowed, secondary CPUs come online through smp_init(), sched_init_smp() finishes the scheduler, workqueue and async machinery start, then do_basic_setup() runs. That function is four lines, and it is where your drivers come up:

static void __init do_basic_setup(void)
{
	cpuset_init_smp();
	driver_init();
	init_irq_proc();
	do_ctors();
	do_initcalls();
}

do_initcalls() walks eight levels in a fixed order, named in init/main.c:

static const char *initcall_level_names[] __initdata = {
	"pure",
	"core",
	"postcore",
	"arch",
	"subsys",
	"fs",
	"device",
	"late",
};

Two details are worth having. The _sync variants are not separate levels: the linker script places .initcallNs.init immediately after .initcallN.init inside the same level, so subsys_initcall_sync() simply runs after every subsys_initcall(). And rootfs_initcall() is not a numbered level at all — it gets its own section, which INIT_CALLS in include/asm-generic/vmlinux.lds.h places between level 5 and level 6. That is how the initramfs is unpacked after filesystems are registered but before device drivers run.

Three steps then complete the kernel side:

	wait_for_initramfs();
	console_on_rootfs();

	/*
	 * check if there is an early userspace init.  If yes, let it do all
	 * the work
	 */
	if (init_eaccess(ramdisk_execute_command) != 0) {
		ramdisk_execute_command = NULL;
		prepare_namespace();
	}

console_on_rootfs() opens /dev/console and dups it onto descriptors 0, 1 and 2, which is how the first userspace program inherits a console it never opened. If it cannot, the kernel prints Warning: unable to open an initial console. and carries on — so that warning means init will start with no standard input or output, which looks exactly like init hanging.

The if is the choice between an initramfs boot and a real root device. ramdisk_execute_command defaults to /init and is overridden by rdinit=. If that path is not accessible the kernel clears it and calls prepare_namespace(), which waits for device probing, parses root=, mounts the root filesystem, mounts devtmpfs and chroots into it. Its failure path prints the available partitions before panicking:

		printk("VFS: Cannot open root device \"%s\" or %s: error %d\n",
				pretty_name, b, err);
		printk("Please append a correct \"root=\" boot option; here are the available partitions:\n");
		printk_all_partitions();
		...
		panic("VFS: Unable to mount root fs on %s", b);

On success you get the single most useful line in this whole region:

	printk(KERN_INFO
	       "VFS: Mounted root (%s filesystem)%s on device %u:%u.\n",
	       s->s_type->name,
	       sb_rdonly(s) ? " readonly" : "",
	       MAJOR(ROOT_DEV), MINOR(ROOT_DEV));

The fixed order in which the kernel searches for an init binary

Back in kernel_init(), after kernel_init_freeable() returns, the kernel frees its init memory, marks rodata read-only, sets system_state = SYSTEM_RUNNING and runs the search. This is the last step before userspace and the order is not negotiable:

	if (ramdisk_execute_command) {
		ret = run_init_process(ramdisk_execute_command);
		if (!ret)
			return 0;
		pr_err("Failed to execute %s (error %d)\n",
		       ramdisk_execute_command, ret);
	}

	if (execute_command) {
		ret = run_init_process(execute_command);
		if (!ret)
			return 0;
		panic("Requested init %s failed (error %d).",
		      execute_command, ret);
	}

	if (CONFIG_DEFAULT_INIT[0] != '\0') {
		ret = run_init_process(CONFIG_DEFAULT_INIT);
		if (ret)
			pr_err("Default init %s failed (error %d)\n",
			       CONFIG_DEFAULT_INIT, ret);
		else
			return 0;
	}

	if (!try_to_run_init_process("/sbin/init") ||
	    !try_to_run_init_process("/etc/init") ||
	    !try_to_run_init_process("/bin/init") ||
	    !try_to_run_init_process("/bin/sh"))
		return 0;

	panic("No working init found.  Try passing init= option to kernel. "
	      "See Linux Documentation/admin-guide/init.rst for guidance.");

Five attempts in order: the initramfs command (/init, or whatever rdinit= set), then init=, then CONFIG_DEFAULT_INIT if it is not empty, then the four hardcoded paths, then panic. Note the asymmetry, because it explains a lot of confusing boots: a failure of init= is a panic() naming the binary and the error number, while a failure of /sbin/init falls through to the next candidate.

How quietly it falls through is the part worth knowing:

static int try_to_run_init_process(const char *init_filename)
{
	int ret;

	ret = run_init_process(init_filename);

	if (ret && ret != -ENOENT) {
		pr_err("Starting init: %s exists but couldn't execute it (error %d)\n",
		       init_filename, ret);
	}

	return ret;
}

On -ENOENT, nothing is printed at all. A board that panics with No working init found may have tried /sbin/init and failed without saying so. Every successful attempt, by contrast, is announced by run_init_process() with Run %s as init process, so the presence or absence of that line is a clean test.

Locating your stage from the console log

These run in order and each one rules something out. Where you need to go below the log, this is the kernel and driver internals the checks rest on.

Check 1 — is anything being buffered that you cannot see? If the console is silent after the bootloader, add an early console before concluding the image is wrong. do_early_param() contains a convenience worth knowing: the console parameter also runs earlycon‘s handler, so a bare console= can bring up an early console on a driver that registers one.

raghu@techveda.org:~$ fw_setenv bootargs "console=ttyS0,115200 earlycon ignore_loglevel root=/dev/mmcblk0p2 rootwait"

If output appears with earlycon and not without it, the kernel was dying before console_init() and you now have the log. If there is still nothing, the kernel is not executing at all, and the problem is the load address, the exception level or the device tree rather than the startup path.

Check 2 — did the command line arrive as you intended? The second line the kernel prints is saved_command_line, the untouched copy. The expected shape, not a capture from a specific board:

raghu@techveda.org:~$ dmesg | grep -m1 "Kernel command line"
[    0.000000] Kernel command line: console=ttyS0,115200 root=/dev/mmcblk0p2 rootwait
raghu@techveda.org:~$ dmesg | grep "Unknown kernel command line parameters"

An empty result from the second command is the pass. A hit names parameters the kernel did not consume and handed to userspace instead.

Check 3 — which initcall did not return? If the log stops partway through driver bring-up, turn on initcall tracing. Its messages are KERN_DEBUG, so ignore_loglevel is needed with them:

raghu@techveda.org:~$ fw_setenv bootargs "console=ttyS0,115200 initcall_debug ignore_loglevel root=/dev/mmcblk0p2 rootwait"

Each initcall then prints calling <function> @ <pid> before it runs and initcall <function> returned <ret> after <usecs> usecs after. The last calling line with no matching initcall line names the function that hung. To boot past it while you investigate, skip it by name — this needs CONFIG_KALLSYMS:

raghu@techveda.org:~$ fw_setenv bootargs "console=ttyS0,115200 initcall_blacklist=my_broken_driver_init root=/dev/mmcblk0p2 rootwait"

Check 4 — did the root filesystem mount? One line settles it, and it also names the filesystem driver that won and whether the mount is read-only:

raghu@techveda.org:~$ dmesg | grep "VFS: Mounted root"
[    2.431002] VFS: Mounted root (ext4 filesystem) readonly on device 179:2.

If that line is absent and you have Cannot open root device followed by a partition list instead, the failure is before init and is a root=, driver or probe-timing question. rootwait covers the probe-timing case.

Check 5 — did the execve succeed, and on what? If the root filesystem mounted and the board still panicked, the failure is the execve, and the file existing is not sufficient evidence. Because try_to_run_init_process() stays silent on -ENOENT, this is the case the kernel will not report, so check it by hand:

raghu@techveda.org:~$ readelf -l /mnt/rootfs/sbin/init | grep interpreter
      [Requesting program interpreter: /lib/ld-linux-aarch64.so.1]
raghu@techveda.org:~$ ls -l /mnt/rootfs/lib/ld-linux-aarch64.so.1
ls: cannot access '/mnt/rootfs/lib/ld-linux-aarch64.so.1': No such file or directory

That output is the expected form of a rootfs missing its dynamic loader, not a capture.

Wrong turns that follow from an incomplete model

Each of these is correct reasoning from a model missing one fact, which is exactly what makes it attractive.

“The console is silent, so the kernel image or the load address is wrong.” Attractive because for most software, no output means it did not start. The missing fact is that console_init() sits near the end of start_kernel(), after setup_arch() and mm_core_init(). A kernel running perfectly well and dying in architecture setup produces exactly the same silence as one that never started. Only an early console separates them, which is why check 1 comes before rebuilding the image twice.

“error -2 means the file is not there, so the root filesystem did not mount.” Attractive because ENOENT does mean “no such file or directory”, and the obvious missing file is init. The missing fact is that execve on a dynamically linked binary also returns -ENOENT when the ELF interpreter named in its program headers is absent. So /sbin/init can be present, correct and executable and the attempt still fails with error 2, because /lib/ld-linux-aarch64.so.1 was never installed into the image. Worse, because the error is -ENOENT, nothing is printed, and the board reports only No working init found. A statically linked init, or readelf -l on the binary, separates the two in one step.

“init= is the parameter to use when something is wrong with init.” Attractive because init= is the documented override and the panic message recommends it. The missing fact is the search order: when an initramfs is in play, ramdisk_execute_command is tried first and defaults to /init, so init=/bin/sh is reached only if /init is absent or fails. The parameter that overrides the initramfs case is rdinit=. Passing init= and seeing no change is this, not a parsing problem.

“Freeing unused kernel memory appeared, so the kernel is done and this is a userspace problem.” Attractive and nearly right, which is why it is worth stating precisely. That line comes from free_initmem() inside kernel_init(), after kernel_init_freeable() has returned and after async_synchronize_full(). It does prove something specific and useful: the initcalls completed, and on a non-initramfs boot prepare_namespace() already mounted the root filesystem. What it does not prove is that any userspace code ran, because it is printed before the first kernel_execve(). Seeing it narrows the problem to the execve, which is a far smaller search than “userspace”.

“Something killed init.” Attractive when the board panics with Attempted to kill init!, because the wording points outward. The missing fact is in kernel/exit.c: that panic fires when the last thread of global init exits, for any reason, including a clean exit with status 0. A shell-script init that runs to completion, or a /bin/sh that reaches end of file on its console, reaches this panic by exiting normally. The exitcode in the message is the clue; a small value usually means init returned rather than crashed.

Key takeaways

  • The chain is primary_entry → start_kernel() → rest_init() → kernel_init() → kernel_execve(), and every console line in between belongs to exactly one of those.
  • console_init() is late in start_kernel(). Silence before it is a logging problem, not necessarily a boot problem, and earlycon is how you tell.
  • PID 1 is created by user_mode_thread(kernel_init, ...) and blocks immediately on kthreadd_done, so PID 2 runs before PID 1 does anything.
  • The boot context becomes PID 0, the idle task, when rest_init() calls cpu_startup_entry().
  • Initcalls run in eight named levels, with rootfs_initcall() in its own section between the fs and device levels, and initcall_debug names the one that hung.
  • The init search order is rdinit= or /init, then init=, then CONFIG_DEFAULT_INIT, then /sbin/init, /etc/init, /bin/init, /bin/sh, then panic.
  • An -ENOENT from the hardcoded candidates is silent, so No working init found does not mean the kernel never tried.

Verification checklist

Run these in order on a board you believe is fixed. Each is an observable, not a reminder.

  1. dmesg | grep -m1 "Kernel command line" prints the command line you intended, character for character.
  2. dmesg | grep "Unknown kernel command line parameters" returns nothing, or names only parameters meant for userspace.
  3. dmesg | grep "VFS: Mounted root" prints one line, with the filesystem type, device and read-only state you expected.
  4. dmesg | grep "unable to open an initial console" returns nothing.
  5. dmesg | grep "Freeing unused kernel memory" prints one line.
  6. dmesg | grep "as init process" names the binary you intended, and no earlier candidate.
  7. cat /proc/1/comm matches that binary.
  8. dmesg | grep "returned with" returns nothing: no initcall left interrupts disabled or returned with a preemption imbalance.
  9. With initcall_debug ignore_loglevel set, every calling line has a matching initcall ... returned line.
Was this worth your time?

Frequently asked questions

What is the first C function the kernel runs on arm64?
start_kernel(). It is called by __primary_switched in arch/arm64/kernel/head.S, after the MMU has been enabled, the stack has been set to init_task‘s and the device tree pointer has been saved to __fdt_pointer.

Why does my board print nothing even though the kernel is running?
console_init() is called near the end of start_kernel(). Everything printed before it, including the banner and the command line, sits in the log buffer until a console is registered. Add earlycon to the command line to see that region.

In what order does the kernel look for init?
The initramfs command first, which is /init unless rdinit= changed it; then init=; then CONFIG_DEFAULT_INIT if it is set; then /sbin/init, /etc/init, /bin/init and /bin/sh in that order; then it panics with No working init found.

Why does the kernel say no working init was found when /sbin/init exists?
Because execve returns -ENOENT when the binary’s ELF interpreter is missing, not only when the binary itself is missing, and try_to_run_init_process() prints nothing for -ENOENT. Check the interpreter with readelf -l and confirm it is present in the image.

What does PID 0 do after the kernel has booted?
It is the idle task on the boot CPU. The context that ran start_kernel() and rest_init() becomes PID 0 when rest_init() calls cpu_startup_entry(CPUHP_ONLINE) and never returns.

Further reading

RB
Raghu Bharadwaj

Founder, TECH VEDA — 20+ years teaching the Linux kernel, device drivers and embedded systems.

Follow on LinkedIn

Get new posts by email

Kernel, embedded Linux and AI-era engineering — a few sharp reads a month. No spam.

We email occasionally and never share your address.