RORK LABJP
BUILD — Rork Max runs real Macs in the cloud loaded with Xcode and the iOS SDK, writing SwiftUI, compiling, reading the errors and building again. That loop, not the code generation, is what lifts the outputNATIVE — What comes out is pure Swift and SwiftUI, not React Native. Reaching AR, Metal graphics and widgets that React Native cannot touch is the real gap between this and other buildersPLATFORMS — Coverage spans iPhone, iPad, Apple Watch, Apple TV and Vision Pro, plus iMessage. Worth a look if you want to start from a watch app or an extension rather than a phone screenCOMPANION — The Rork Companion app lets you check a generated build on a real iPhone without a paid Apple Developer account, lowering the bar for trying a first project end to endPRICING — Free to start, paid plans from $25 a month, and Rork Max on the $200 Max plan. Worth working out up front how many projects it takes to earn that backDEADLINE — From August 31, 2026, Google Play requires target API level 36 or higher for new apps and updates alike. Ten days out, and the targetSdkVersion of what you generate is yours to verifyBUILD — Rork Max runs real Macs in the cloud loaded with Xcode and the iOS SDK, writing SwiftUI, compiling, reading the errors and building again. That loop, not the code generation, is what lifts the outputNATIVE — What comes out is pure Swift and SwiftUI, not React Native. Reaching AR, Metal graphics and widgets that React Native cannot touch is the real gap between this and other buildersPLATFORMS — Coverage spans iPhone, iPad, Apple Watch, Apple TV and Vision Pro, plus iMessage. Worth a look if you want to start from a watch app or an extension rather than a phone screenCOMPANION — The Rork Companion app lets you check a generated build on a real iPhone without a paid Apple Developer account, lowering the bar for trying a first project end to endPRICING — Free to start, paid plans from $25 a month, and Rork Max on the $200 Max plan. Worth working out up front how many projects it takes to earn that backDEADLINE — From August 31, 2026, Google Play requires target API level 36 or higher for new apps and updates alike. Ten days out, and the targetSdkVersion of what you generate is yours to verify
Articles/App Dev
App Dev/2026-06-28Advanced

Design On-Device Core ML So Cold Start and Heat Don't Break It

Put on-device Core ML in the native Swift that Rork Max generates and you hit two walls before accuracy: the first inference is slow, and the device heats up and slows down. Here is a design built around cold start and a thermal budget, with working Swift.

Rork Max233Core ML6Swift48on-device AI5indie developer39

Premium Article

Because Rork Max can generate native Swift apps, on-device Core ML inference — long out of easy reach in React Native — is now within reach even for indie development. But put it on a real device and there are stumbling points before you ever get to accuracy. The first inference is oddly slow. After a while the device warms up and everything feels sluggish. Both are design problems about when and how much you run inference, not about whether the model is good.

As an indie developer at Dolice, when I built on-device inference into an app I run, I struggled with a roughly one-second freeze on the very first call. The cause was not model accuracy; it was running the first load and inference on the main thread right at launch.

This article lays out how to design Core ML around two constraints — cold start and a thermal budget — with Swift code.

Why the first inference is slow

A Core ML model runs two heavy operations the first time you use it. One is loading and compiling the model (optimized for the device's Neural Engine); the other is the first inference, which allocates internal buffers. It is normal for the first call to be an order of magnitude slower than later ones.

StageWhat mainly happensFelt impact
First loadModel compile and placementHundreds of ms to 1 s
First inferenceBuffer allocation, warmupTens to hundreds of ms
SubsequentRun on the allocated pathOften around 10-30 ms

The problem is throwing that heavy first call at the moment the user is waiting for a result. The design goal is simple: move the heavy first call earlier, to a time when the user is not waiting.

Pull model load and warmup out of the launch flow

First, do not synchronously load the model at app launch. Defer with lazy, and do the load and warmup (one inference on dummy input) on a background queue.

import CoreML
 
actor InferenceEngine {
    private var model: MyModel?
 
    // warm up in the background; call after the first screen appears, not at launch
    func warmUp() async {
        guard model == nil else { return }
        let config = MLModelConfiguration()
        config.computeUnits = .all        // let it use the Neural Engine too
        do {
            let loaded = try MyModel(configuration: config)
            // run the first inference on dummy input to allocate buffers
            _ = try? loaded.prediction(input: .dummy)
            model = loaded
        } catch {
            model = nil                   // do not block launch even on failure
        }
    }
 
    func predict(_ input: MyModelInput) async throws -> MyModelOutput {
        if model == nil { await warmUp() }
        guard let model else { throw InferenceError.unavailable }
        return try model.prediction(input: input)
    }
}

The caller invokes warmUp() in the brief gap after the first screen renders and before the user starts interacting.

.task {
    // warm up during idle time after the screen shows
    await engine.warmUp()
}

The actor is what pays off here. Even if predict is called from several places at once, the language guarantees the load does not run twice. Early on I forgot this protection and multi-loaded the model on every screen transition, needlessly bloating memory. Concurrent access is a quiet pitfall in on-device inference.

Thank you for reading this far.

Continue Reading

What follows includes implementation code, benchmarks, and practical content we hope you'll find useful. This site runs without ads — server and development costs are supported entirely by members like you. If it's been helpful, we'd be truly grateful for your support.

WHAT YOU'LL LEARN
What cold start really is (first compile, first inference) and pulling it out of the launch flow with an actor
Gating that reads ProcessInfo.thermalState and steps inference down across full / reduced / suspended
Releasing the model on memory warnings and backgrounding, then warming it up again on return
Secure payment via Stripe · Cancel anytime

Unlock This Article

Get full access to the rest of this article. Buy once, read anytime. This site is ad-free — your support goes directly toward keeping it running.

or
Unlock all articles with Membership →
Share

Thank You for Reading

Rork Lab is ad-free, supported entirely by members like you. We publish practical guides daily with implementation code, benchmarks, and production-ready patterns. If you've found it useful, we'd love to have you on board.

  • Copy-paste ready implementation code
  • New advanced guides published daily
  • $5/mo or $10 for lifetime access
View Membership →

Related Articles

App Dev2026-07-05
Rork Max (Swift) or the Standard Version (React Native): How to Decide as a Solo Developer
Stuck between Rork Max's native Swift and the standard React Native version? Here is a practical decision framework built from a solo developer's perspective, weighing cost, feature boundaries, and how easy each path is to migrate later.
App Dev2026-07-03
Making Your Rork Max App Resilient to Dropped and Restored Connections: Offline Detection and Retry with NWPathMonitor
Build networking that survives a lost signal in your Rork Max native Swift app with NWPathMonitor. Detect offline states, respect Low Data Mode and cellular, and auto-resend queued work on reconnect — all with working Swift code.
App Dev2026-07-03
Keeping Downloads Alive After Your Rork Max App Is Killed: Background URLSession Design and Relaunch Handling
How to design downloads in a Rork Max native Swift app so transfers continue in the OS daemon even after the app is suspended or terminated. Covers relaunch wiring, resumeData recovery, and measured isDiscretionary behavior with working code.
📚RECOMMENDED BOOKS
Build a Large Language Model (From Scratch)
Sebastian Raschka
LLM Dev
Prompt Engineering for LLMs
Berryman & Ziegler
Prompting
AI Engineering
Chip Huyen
AI Eng
* Contains affiliate links
See all →