ENAS runs NAS on CPUs and slashes search time
Neural architecture search usually eats up GPU hours before you even deploy a model to a microcontroller. A new framework called ENAS skips the GPU requirement entirely. It uses a hardware-aware approach that fits tight SRAM limits on tiny chips. The core idea is simple but effective: check feasibility early, search within a constrained space, and cache results aggressively. This matters if you are building TinyML models for devices with less than a megabyte of memory.
The framework operates without GPU acceleration. That alone makes it viable for resource-constrained dev environments where renting compute isn't practical. ENAS relies on three main components. First is a static feasibility check. Second is a cell-based search space. It supports standard blocks, depthwise-separable convolutions, and bottleneck blocks. Skip connections are optional. Third is a three-stage hybrid search strategy. It starts with random sampling, moves to top-K selection, and finishes with mutation. Cross-run caching persists between sessions to avoid redundant calculations.
I looked at the benchmarks to see how it performs against existing tools like NanoNAS. The authors tested on Visual Wake Words and Melanoma Cancer datasets. They deployed across eight different microcontrollers. These chips had SRAM ranging from 20 KB to 1 MB. Nine input image resolutions were also tested. The speedup numbers are significant for CPU-only workloads. On Visual Wake Words, ENAS achieved a 2.41x mean search-time speedup. For Melanoma Cancer, it hit 1.70x. Test accuracy remained competitive with NanoNAS despite the lighter computational footprint.
The real win is memory efficiency. Microcontroller deployment often fails because peak activation RAM exceeds available SRAM. ENAS-selected models use substantially lower peak activation RAM at matched accuracy levels. In a direct test on an STM32H743-based microcontroller, ENAS reached 79.4% test accuracy. This outperformed the greedy CPU-only baseline by 2.6 percentage points. That margin is critical when every kilobyte counts.
Setting up the workflow is straightforward since there is no CUDA dependency. You define the target hardware constraints first. ENAS filters architectures that won't fit before they are fully evaluated. This static check prevents wasted cycles on invalid candidates. The cell-based design keeps the search space manageable. Depthwise-separable and bottleneck blocks are efficient primitives for edge devices. Optional skip connections allow the optimizer to find deeper paths without bloating parameters.
The three-stage search strategy balances exploration and exploitation. Random initialization covers broad areas of the search space. Top-K selection focuses resources on promising candidates. Mutation refines those candidates for better performance. Persistent caching ensures that repeated evaluations of similar cells don't cost extra time. If you run multiple experiments, the cache pays off immediately. You aren't paying for the same calculation twice.
Memory footprints vary significantly across microcontrollers. A 20 KB SRAM chip requires aggressive pruning or very shallow networks. A 1 MB chip allows more flexibility. ENAS adapts to both ends of the spectrum. The framework outputs models tailored to specific hardware limits. This eliminates the guesswork of manual compression. You get an architecture that fits by design rather than by trial and error.
Accuracy remains stable across different resolutions. Input size affects compute load and memory usage. ENAS maintains performance even as resolution changes. This consistency is useful for product scaling. You can target multiple devices with similar network structures. The open-source release at GitHub allows immediate testing. Check the EdgeIntelligenceLab repo for implementation details.
For TinyML projects, GPU-free NAS removes a major barrier. ENAS delivers speedups without sacrificing accuracy. Lower peak RAM usage solves a common deployment failure mode. If you are constrained by hardware budget or compute access, this framework offers a practical alternative. The STM32H743 results prove it works on real silicon.
I once spent weeks on GPU-based NAS only to have my model fail on a device with less than a megabyte of memory. ENAS running on CPUs feels like a sarcastic victory for my sanity.