← Writing

Engineering· Updated

On embedded, the worst bugs are often not in your code

A thread of real war stories from r/embedded shows a pattern: the bugs that cost the most time are rarely logic errors. Why embedded debugging is mostly search, and how firmware observability makes the search shorter.

This month a thread hit the top of r/embedded with a simple prompt: share the embedded bug that cost you hours. It drew hundreds of replies. I read through about ninety of them looking for a pattern, and the pattern surprised me.

Almost none of the bugs were logic errors.

That thread asked for the most memorable bugs, so it is not a random sample. Most day-to-day firmware bugs are still ordinary code bugs. But the ones that eat days tend to look like these.

Three real stories

One engineer spent a day hunting ten or twenty microamps of current draw that would not go away. The reading was maddeningly random. At one point a colleague walked over to help and the extra current vanished; they stepped away and it came back. The cause was sunlight from a window landing on a chip with an exposed die, generating a small photocurrent. A shadow fixed it. Several other people replied with the same class of story, including an ASIC that refused to power up until someone shone a flashlight on the die.

Another described a board that stopped transmitting serial data whenever the air got humid. Breathe on it and it would go quiet. They suspected the humidity sensor and pulled it out entirely. The board still failed. The real cause was a cold solder joint on the enable line of the RS-232 transceiver. With the joint barely connected, the pin was effectively floating, and in humid air the leakage across the board surface pulled it to the wrong level.

A third shipped a batch of devices with the wrong crystal, a few percent fast. Everything worked except a routine that detected a tone, because that was the one thing that depended on the exact sample rate. Weeks went into tracing it, because the crystal was nowhere near where the symptom showed up.

A separate post the same month came from someone so tired of chasing I²C problems with an oscilloscope that they built a dedicated diagnostic tool. The comments turned into a catalogue of the ways a two-wire bus lies to you: a target holds the clock low longer than the controller tolerates and the transaction fails as if the device were not there; the processor resets mid-byte and leaves a target holding the data line low until someone clocks it free or power-cycles it. As one commenter put it, when the data line is stuck low the fix is completely different depending on which chip is holding it, and right now everyone guesses. Another summed the bus up: I²C is barely a standard. The spec exists; the trouble is how loosely devices follow it.

The pattern

Read enough of these and the shape is always the same. The fix, when it finally arrives, is trivial. A shadow. A reflow. The right crystal. The cost was never the fix. It was the search. And the search was long because the fault was not where the symptom showed up.

This is what makes embedded debugging feel different, and why it humbles people who are strong software engineers. In application software, most bugs are in logic you can inspect. You can read it, log it, step through it, reproduce it. Your whole toolkit assumes the fault lives somewhere you can look.

Embedded systems break that assumption. The system is not just your code. It is your code plus a physical world your code cannot see: light, heat, humidity, a marginal solder joint, a clock a few percent off, a bus that half the devices on it implement differently. The symptom appears in software, so that is where you start looking, but the cause is usually a layer or two below, in a place your instincts never point you.

That gap, between where the symptom shows and where the cause lives, is the entire debugging cost.

What actually helps

You cannot make the physical world stop misbehaving. What you can do is shrink the gap, by making the invisible parts of the system visible. That is really what firmware observability is: not a product category, but the practice of getting the system to tell you what it is doing so you spend less time guessing.

A few things move the needle, roughly in the order they start helping:

  • See the physical layer directly. A logic analyzer on the bus turns “the sensor is flaky” into decoded transactions you can read, and a scope shows the analog faults a logic analyzer hides, like slow edges from weak pull-ups. The person who built that I²C tool was doing exactly this, just packaged so they never had to set up the scope again.
  • Capture the crash instead of losing it. In the field, a fault usually ends in a reset, from a watchdog or a reset-on-fault handler, and the RAM state that would explain it is gone. A fault handler that saves a coredump to flash means the one crash you did catch still tells you something after the reset. On Zephyr, the coredump subsystem writing to a flash partition, plus GDB on the host, gets you a backtrace after the device has rebooted, as long as you kept the zephyr.elf from the build that crashed. On the bench, a probe channel like SEGGER RTT lets you watch the device live without a spare UART.
  • Get logs off the device without changing the timing. The classic embedded trap is that adding a print statement moves the bug. Loggers that move formatting off the device, such as Rust’s defmt or Zephyr’s dictionary-based logging, send references to the format strings and the raw arguments, and let the host build the strings, which keeps the on-device cost small. The host can only do that with a file from the exact build running on the device: the ELF for defmt, and for Zephyr the log_dictionary.json that the build writes next to the image. Zephyr’s logging documentation says the dictionary only works with the build that produced it, so archive it with every image you ship; rebuilding the same source later is not a safe substitute. It also pays to have each device log an identifier for its build, for example at boot, so a log from the field points straight to the right archived files. Trace tools like SEGGER SystemView or Percepio Tracealyzer show you the ordering of interrupts and tasks that a print can never capture.
  • For the bugs that only happen in the field, bring the field to you. The hardest stories above share a trait: they showed up on one unit, in one place, under conditions you cannot reproduce on the bench. This is where firmware observability platforms earn their place. They ship crash reports, metrics, and logs home from deployed hardware, so you can see the state that led to a fault without standing next to the device. Memfault collects coredumps, metrics, and logs, and adds OTA updates. Golioth is a broader device platform that covers remote logs, diagnostics, and OTA, and hands crash analysis to Memfault through an integration. Spotflow, where I work, collects logs, metrics, and core dumps with stack traces, with ready-made integrations for Zephyr, ESP-IDF, and Nordic’s nRF Connect SDK. It sends data over MQTT and buffers it on the device while the connection is down, which suits devices that are only online now and then, and it also handles OTA updates and tracks known CVEs in your firmware. I am telling you that I work there so you can weigh the recommendation accordingly.

None of this stops the wrong-crystal bug from happening. What it changes is how long you spend in the dark before you find it.

The real skill

So here is the thing I would tell someone new to embedded, and the reason that thread stuck with me. The engineers who are fast at this are not the ones who write bug-free firmware. Nobody does. They are the ones who have learned, usually the hard way, that the bug is probably not where they are looking, and who have built enough visibility into the system that the search is short.

You cannot add that visibility at three in the morning, with a customer’s device failing in a building you cannot get to. You add it before, on purpose, while everything still works. That is the discipline. The debugging skill is really a seeing skill.

Updated 4 October 2026: thanks to @ManBoEmbedded on X for pointing out that the logging dictionary has to be archived with every build, and that logs should name the build they came from. The logging and coredump advice above now covers both.

Esc