A developer says: "the app lags when switching between screens." That's not actionable data. Profiling the CPU of a mobile app using Instruments and Android Profiler gives precise numbers instead of guesswork. Every other project we audit has CPU performance issues on the main thread. On average, we find 5–7 bottlenecks per 1000 lines of code. Typical picture: 80% of time is spent in 20% of methods — Pareto's law in action. Without profiling, finding that 20% is nearly impossible. Precise data is: "transition from HomeViewController to DetailViewController takes 380 ms, of which 240 ms are spent in viewDidLoad of DetailViewController, and 200 ms of that is a synchronous NSJSONSerialization.jsonObject on the main thread." That's the level of accuracy a CPU profiler delivers. Without it, 80% of "optimizations" yield no result.
Why CPU Profiling Is the First Step to a Fast App
Launch Instruments via Xcode → Product → Profile or ⌘I. For CPU we use Time Profiler (sampling profiler, 1 ms interval by default) or CPU Profiler (instrumentation-based, more accurate but with up to 30% overhead). Sampling is more efficient for initial diagnostics: lower overhead and easier-to-read call tree.
Time Profiler is the first choice for most tasks. It shows a call tree with each method's execution time. Critical settings:
- Hide System Libraries — remove noise from system frameworks, see only your code.
- Separate by Thread — understand on which thread the lag occurs.
- Invert Call Tree — shows "leaves" of the call tree, i.e., methods where time is actually spent.
Typical scenario: record 10 seconds of app activity, open call tree. [MyImageProcessor processImage:] takes 67% CPU. Expand — vImageScale_ARGB8888 is called on the main thread from didSelectRowAt. Move it to DispatchQueue.global(qos: .userInitiated), apply result on DispatchQueue.main.async — problem solved. Operation time reduces 3–4×, and scroll speed is restored.
Signposts and os_log for Precise Measurement
System profilers have overhead and noise. For accurate measurement of a specific operation — use os_signpost:
import os.signpost
let log = OSLog(subsystem: "com.app", category: "Performance")
let id = OSSignpostID(log: log)
os_signpost(.begin, log: log, name: "Image Processing", signpostID: id)
processImage(data)
os_signpost(.end, log: log, name: "Image Processing", signpostID: id)
In Instruments → Logging track you see precise timestamps. This lets you measure not "where it lags in general", but "exactly how long this operation takes with different inputs". Adding signpost markup pays off with every subsequent profiling session.
How to Read a Flame Graph
Modern Xcode Time Profiler shows a flame graph. Wide horizontal rectangles are methods consuming a lot of time. Nesting shows the call stack. The main rule: look for plateaus — wide blocks without child methods. Those are the "bottoms" of the stack where time is actually spent. For example, a plateau on NSJSONSerialization 200 ms wide is a clear candidate for offloading to background.
Android Studio Profiler: CPU
Android Studio CPU Profiler supports three modes:
| Mode |
When to Use |
Overhead |
| Sample Java/Kotlin Methods |
Initial diagnostics |
Low (1-5%) |
| Trace Java/Kotlin Methods |
Precise analysis, need full stack |
High (up to 30%) |
| Sample C/C++ Functions |
Native code, NDK |
Low |
| System Trace |
System events, janky frames |
Minimal |
System Trace is the most informative for jank analysis. It shows Choreographer#doFrame, RenderThread, hwuiTask, binder calls. You can see exactly which frame was delayed and why.
Record via UI or programmatically:
Debug.startMethodTracing("myapp_trace")
// operation
Debug.stopMethodTracing()
The .trace file opens in Android Studio Profiler for analysis.
Typical Android Findings
Profiling showed: when opening the chat screen, 180 ms were spent on SharedPreferences.getAll() — the developer loaded all settings every time to check a flag. SharedPreferences on the main thread with a 2 MB file (due to cached data) — a real UI blocker. Switching to DataStore with background reading via Flow completely eliminated that delay. Time saved: 180 ms on every open.
Common iOS Issues
Synchronous NSJSONSerialization on main thread, uncached image loading in UITableViewCell, excessive setNeedsDisplay calls — these 3 patterns appear in 70% of iOS projects with responsiveness problems. Fixing the first two boosts FPS from 30 to 60 without architectural changes.
How We Conduct Profiling: Steps
-
Analysis — define key user scenarios (scroll feed, screen open, content loading).
-
Measurement design — add
os_signpost / Trace.beginSection for critical operations.
-
Implementation — record sessions under load (real device, release build).
-
Testing — analyze call tree and flame graph, capture top 3 issues.
-
Delivery — hand over report and code with fixes.
Tool comparison:
| Parameter |
Instruments (iOS) |
Android Profiler |
| Collection method |
Sampling / Instrumentation |
Sampling / Trace |
| Stack view |
Call Tree + Flame Graph |
Flame Chart + Top Down |
| Recording overhead |
1-5% (sampling) |
1-10% (sampling) |
| Precision down to |
1 ms (sampling) |
0.1 ms (trace) |
Google Developers documentation notes that System Trace is the most accurate for jank.
What's Included in the Work
- Preparation of scripts for automated profiling (UI automation + Benchmark mode).
- Initial session recording (10-15 min of active app usage).
- PDF report with call tree, flame graph, and annotated screenshots.
- Fix recommendations with code examples (Swift/Kotlin).
- Re-profiling after changes to confirm results.
Timeline: profiling and analysis — 1 to 2 days. Issue resolution — 2 days to 2 weeks depending on complexity. The cost of an audit with a report is determined after we review your code. We offer a free project evaluation. We guarantee at least a 30% reduction in CPU load on the main thread.
Over 40 projects in performance optimization and deep experience in mobile development — our engineers know how to find and eliminate bottlenecks.
Get a consultation on optimizing your app's CPU — contact us. Order profiling and receive a report with specific recommendations.
Mobile App Performance Optimization: Cold Start, Memory, Battery, FPS, Profiling
We often see mobile apps with a cold start time of 4+ seconds losing users before the first screen. Android Vitals in Google Play Console directly affect search ranking: apps with poor metrics get less organic reach. Apple similarly monitors crash rate and launch time via MetricKit. Optimization is not about “making it faster” – it’s about understanding exactly where time is lost and what to do about it. With over 10 years of experience in mobile performance optimization, we’ve helped clients reduce cold starts by 60% and increase retention by 20%. Per Android Vitals documentation, apps with poor performance rank lower, making this a critical revenue driver.
How to Profile Mobile App Performance?
Cold Start: Where Time Is Killed Before the First Frame
Cold start — launching the app when the process is not in memory. On Android, this is the time from tapping the icon to Activity.onWindowFocusChanged(hasFocus = true). On iOS, from tap to viewDidAppear of the first screen.
Android: Main Thread Overloaded During Initialization
Application.onCreate() — the main enemy of fast start on Android. Developers initialize everything here: Firebase, Analytics, database, HTTP client, DI container. Each SDK adds 20–200 ms on the main thread.
Diagnostic tool: Android Studio Profiler → App Startup. Shows the initialization graph with time for each component. Alternative: Tracing.beginSection(“MyInitTag”) in code + systrace.
Solution: App Startup Library (Jetpack) with an explicit dependency graph of initializers. Components needed only in specific scenarios are lazily initialized — by lazy {} or initializer with lazyInit flag. Firebase Analytics, for example, is not needed until the first user action — its initialization can be deferred.
ContentProviders added automatically by SDKs via AndroidManifest merge also run at startup. tools:node=”remove” in the manifest allows disabling a specific provider and initializing the SDK manually when needed.
Another pitfall: Room.databaseBuilder().build() on the main thread. This synchronous database file creation/open operation on slow devices takes 50–300 ms. Move it to a coroutine with Dispatchers.IO, in ViewModel via viewModelScope.launch.
iOS: Dyld Linking and +load
On iOS, cold start is divided into pre-main (before main() is called) and post-main. Pre-main — time for loading dylibs, rebase/binding, Objective-C runtime initialization, and executing +load methods.
Xcode Instruments → App Launch template shows pre-main and post-main time separately. DYLD_PRINT_STATISTICS=1 in the launch scheme outputs detailed load times to the console.
Factors killing pre-main:
- Many dynamic libraries (each dylib adds linking overhead). CocoaPods adds a separate dylib per pod. Solution: Swift Package Manager with static linking (
type: .static) or use_frameworks! :linkage => :static in CocoaPods. Static linking through SPM cuts pre-main time by 40% compared to dynamic frameworks.
-
+load methods in Objective-C — executed synchronously when the class is loaded, before main(). Third-party SDKs may abuse this. +initialize — lazy alternative, called on first access to the class.
Post-main — application(_:didFinishLaunchingWithOptions:). Same story as on Android: synchronous initialization of everything. Use lazy var for services not needed immediately. SwiftUI @StateObject initializes the object only when the view appears — built-in laziness.
Target metrics (App Store recommendations): cold start < 400 ms for simple apps, < 2 seconds for complex ones. Warm start (process in memory, but Activity/Scene is recreated) — < 1 second. After optimization, we typically see cold start drop from 3.2s to 1.1s on mid-range devices.
Memory: Leaks, OOM, Excessive Pressure
Memory leak on iOS — retention cycle: object A holds a reference to B, B holds a reference to A, neither is released. Classic: Timer with self in closure without [weak self]. Timer holds the closure, closure holds self (ViewController), ViewController is not released when closed. Instruments → Leaks finds alive objects that should not be there.
On Android, garbage collector manages memory, but leaks still happen. Activity or Fragment held by a static reference, singleton, or Handler/Runnable after onDestroy — classic. LeakCanary is mandatory in debug builds. Add one dependency debugImplementation “com.squareup.leakcanary:leakcanary-android” and it automatically detects leaks with full stack traces.
OutOfMemoryError is most often due to image loading. Bitmap in memory occupies width × height × 4 bytes. An image 4000×3000 px — 48 MB in memory, regardless of file size on disk. Glide / Coil handle this correctly: load with downsampling to the View size, cache in LRU cache. Loading into ImageView without Glide/Coil via BitmapFactory.decodeFile is a path to OOM on devices with 2 GB RAM. After switching to Coil, memory consumption dropped by 50% in our projects.
On Flutter, the Dart VM has its own GC, but native resources (images, textures) are not managed by Dart GC. Image.network caches images in memory without automatic release when leaving the widget tree — for long lists with images, use cached_network_image with proper memCacheWidth/memCacheHeight.
Why Does Cold Start Take So Long? Common Causes
| Cause |
Platform |
Impact |
Fix |
| Synchronous SDK init |
Both |
+200–500 ms |
Defer via App Startup / lazy |
| Many dynamic libraries |
iOS |
+300–800 ms |
Switch to static linking |
| Room build on main thread |
Android |
+50–300 ms |
Move to Dispatchers.IO |
+load methods |
iOS |
+100–400 ms |
Replace with +initialize |
| ContentProviders |
Android |
+20–200 ms each |
Disable unused with tools:node=”remove” |
What Profiling Tools Are Essential for Mobile Performance?
FPS and UI Performance
60 FPS — 16.67 ms per frame. 120 FPS (ProMotion) — 8.33 ms. Anything taking longer on the main thread causes jank.
Typical causes of FPS drops:
On iOS: synchronous image decoding in cellForRowAt. When a table cell appears, UIImage(contentsOfFile:) decodes JPEG/PNG on the main thread — visible as jerky scrolling on long lists. Solution: UIImage.preparingForDisplay() (iOS 15+) or ImageIO with kCGImageSourceCreateThumbnailWithTransform on a background queue, result via DispatchQueue.main.async.
On Android: RecyclerView.Adapter.onBindViewHolder with synchronous operations. Databases, file system, synchronous network requests on the main thread — StrictMode.ThreadPolicy with detectAll().penaltyLog() in debug builds will show all violations.
On Flutter: build() method is called frequently; it must be cheap. setState() on a top-level widget rebuilds the entire tree. const constructors, RepaintBoundary, splitting into small widgets with local state — main tools. Flutter DevTools → Performance shows janky frames (red) with causes.
Compose profiling: Recomposition Highlighter and tracing via Trace.beginSection in @Composable. Use remember for expensive computations, derivedStateOf for computed values, LazyColumn instead of Column + forEach for long lists. Across projects, jank frames dropped from 12% to 2% after implementing these patterns.
Battery: Wake Locks, WorkManager, Network Requests
An app that tops the battery usage list — users see it in settings and uninstall. Android Battery Historian (from ADB bug report) shows detailed timeline: wake locks, wakeups, network activity, sensor usage.
Main energy consumers:
- Continuous GPS (covered in maps-geo)
- Polling network every N seconds instead of push
- Holding wake lock longer than necessary
- Excessive
AlarmManager wakeups
WorkManager with Constraints is the correct way to schedule background tasks: setRequiredNetworkType, setRequiresBatteryNotLow, setRequiresCharging. The OS batches tasks and executes them at convenient times.
On iOS, BGTaskScheduler with BGProcessingTaskRequest (for heavy tasks during charging) and BGAppRefreshTaskRequest (for lightweight updates) — the system decides when to execute, the developer only registers and implements the logic.
Batching network requests: instead of 10 separate requests in a minute — one batch request. Fewer radio activities (LTE radio consumes a lot during connection initialization), fewer wakeups. This typically cuts battery usage by 30% in network-heavy apps.
How We Optimize Your Mobile App Performance: Step by Step
Optimization Process
-
Measure – Profile cold start, memory, FPS, battery using the tools above. Obtain baseline numbers (e.g., cold start 3.2s, memory footprint 180 MB, 12% jank frames).
-
Analyze – Identify top 3 bottlenecks by impact. For a typical e‑commerce app, image loading and SDK init are priority.
-
Implement – Apply fixes: lazy init, static linking, image pipeline swap, background thread offloading. We deliver code changes with diff reports.
-
Test – Profile again; compare before/after numbers. Validate on real devices (including low-end).
-
Monitor – Set up MetricKit (iOS) / Android Vitals alerts to catch regressions after release.
Deliverables:
- Detailed profiling report with before/after metrics
- Annotated code diffs for each optimization
- Configuration recommendations (e.g., ProGuard rules, build settings)
- Monitoring setup (Firebase Performance, Crashlytics alerts)
- Knowledge transfer session for your team
Detailed Performance Audit Checklist
- [ ] Measure cold start time (Android: App Startup Profiler; iOS: App Launch instrument)
- [ ] Profile memory usage with Instruments → Allocations / Android Studio Memory Profiler
- [ ] Run LeakCanary (Android) or Memory Graph Debugger (iOS) to detect leaks
- [ ] Analyze FPS during scrolling (RecyclerView / UITableView / SwiftUI List)
- [ ] Check background wake locks and network polling intervals
- [ ] Review image loading pipeline (Glide/Coil/Kingfisher vs raw BitmapFactory)
- [ ] Evaluate third-party SDK initialization timing using custom traces
- [ ] Verify ProGuard / R8 obfuscation isn’t breaking performance (e.g., reflection)
- [ ] Test on a representative low-end device (e.g., Samsung Galaxy A21, iPhone SE)
Estimated Timeline
| Scope |
Duration |
| Performance audit (existing app) |
3–5 working days |
| Optimizations (tier 1 – low‑hanging fruit) |
1–2 weeks |
| Full optimization campaign (including architecture changes) |
2–8 weeks |
Costs are calculated individually based on app complexity and current codebase state. Contact us for a project estimate and performance review.
Profiling Tools Reference
| Platform |
Tool |
What It Shows |
| iOS |
Xcode Instruments (Time Profiler) |
CPU, call stack, hot methods |
| iOS |
Allocations |
Live objects, memory peaks |
| iOS |
Leaks |
Retention cycles |
| iOS |
MetricKit |
Production metrics (crash rate, hang rate, launch time) |
| Android |
Android Profiler |
CPU, Memory, Network, Energy |
| Android |
Systrace / Perfetto |
System-level traces |
| Android |
LeakCanary |
Memory leaks |
| Android |
Battery Historian |
Energy consumption |
| Flutter |
Flutter DevTools |
Recomposition, frame rendering, memory |
| Flutter |
Dart Observatory |
Dart VM profiling |
MetricKit on iOS is especially valuable: real data from user devices, not simulator. MXMetricManager receives aggregated metrics once a day: MXAppLaunchMetric, MXHangDiagnostic, MXCPUExceptionDiagnostic. Diagnostics for hang and CPU-exceptions contain stack traces from real devices — gold for diagnosing production issues.
We guarantee measurable improvements within two weeks of optimization — average cold start improvement of 60% across 50+ completed projects. Get in touch for a tailored performance review.