We often see how improperly packed textures turn a game into a "heavy weight": 300 MB instead of 80 MB, and on old Android devices it crashes with SIGKILL due to OOM. Meanwhile, everything works fine in the editor — Unity stores textures in uncompressed form, and the real VRAM size becomes apparent only when profiling on a device. Our team, with 5+ years of experience on over 30 projects, has achieved a 50-70% reduction in memory consumption.
Optimizing Texture Atlases: Key Methods
The first and most common cause of memory overuse is the default RGBA32 format. If you don't set Override for each platform, Unity packs textures in uncompressed RGBA32. A 1024×1024 texture in RGBA32 takes 4 MB of VRAM. The same data in ASTC 6×6 is around 250 KB. That's a 16x difference. Always check compression settings for each platform.
The second typical mistake is an atlas size that is not a power of two. The GPU processes 700×900 textures inefficiently. Unity issues a warning but does not fix it automatically. Non-POT textures are not compressed by most formats and take more memory than the nearest POT. We guarantee bringing all atlases to POT.
The third problem is transparency where it's not needed. If the sprite is opaque, use RGB instead of RGBA. This is easily set in the import settings.
Compression Format Comparison Table
| Platform |
Format |
Quality |
Usage |
Memory Savings (vs RGBA32) |
| Android (modern) |
ASTC 6×6 |
high |
all textures |
up to 16x |
| Android (legacy, API < 23) |
ETC2 |
medium |
opaque; ETC2 RGB for alpha |
up to 8x |
| iOS (A8+) |
ASTC 6×6 |
high |
all textures |
up to 16x |
| iOS (fallback) |
PVRTC |
low |
only when memory is extremely limited |
up to 6x |
ASTC is adaptive and supports any aspect ratio. Available on all devices with Android 5.0+ and iOS A8+. For game sprites, 6×6 gives optimal quality/size. For UI with small text, use 4×4. In terms of quality, ASTC 6×6 is 2-3 times better than PVRTC at the same size.
ETC2 on Android is only needed for devices older than five years. If analytics show the share of such devices > 5%, take it into account.
How to Choose the Right Compression Format?
The choice depends on target devices. For most modern projects, ASTC 6×6 is optimal. If the project targets old devices, use ETC2 for Android and PVRTC for old iOS. It's important to set Overrides in the import settings for each platform.
Why Split Atlases by Context?
The correct strategy is to split by usage context, not by object type:
- ui-hud: HUD elements displayed constantly
- ui-menus: buttons, backgrounds, icons for menus (unloaded during gameplay)
- gameplay-player: player sprites and animations
- gameplay-enemies-tier1: first-level enemies
- gameplay-vfx: particles, explosions, effects
Such splitting allows unloading unused atlases via Resources.UnloadUnusedAssets() or Addressables with explicit Release. With a single atlas, this is impossible. One of our clients, after restructuring 12 atlases, reduced peak memory by 35%.
Impact of Empty Space
Unity SpriteAtlas leaves padding between sprites (default 4 px) to prevent bleeding. On a 1024×1024 atlas, area loss can reach 15-20%. Reduce padding to 2 px for sprites without subpixel rendering — this improves packing. Monitor fill rate in the atlas Preview. If less than 70%, rework the split.
Process and Timeframes
Turnkey stages (3-6 days)
- Analytics: profiling on a real device via Memory Profiler, identifying top-10 heavy textures.
- Design: setting Overrides for Android/iOS, restructuring by context, disabling Mip Maps for 2D sprites.
- Implementation: modifying atlases, checking POT sizes, reducing padding.
- Testing: final check using Android GPU Inspector or Xcode Metal System Trace, comparing VRAM before and after.
- Documentation: report with results and maintenance recommendations.
Typical Results Table
| Parameter |
Before Optimization |
After Optimization |
Reduction |
| VRAM for textures |
~300 MB |
~90 MB |
70% |
| Number of textures |
120 |
120 |
unchanged |
| Draw Calls per scene |
~200 |
~150 |
25% |
Figures are for an average project. Your results may vary.
What’s Included
- Full audit of current textures and atlases using Memory Profiler.
- Compression Override setup for each target platform.
- Atlas restructuring by usage context.
- Optimization of padding, Mip Maps, and POT sizes.
- Testing on real devices (Android and iOS).
- Documentation with recommendations.
- Post-optimization support for 1 month, including access to all reports and documentation.
Pricing
Optimization costs depend on the project scale. A basic audit starts at $500, full optimization from $1500. Contact us for a free consultation.
Result
After optimization, you will reduce texture memory usage by 50-70%, decrease Draw Calls due to proper packing, and ensure stable performance on devices with 2-3 GB of RAM. Visual quality will remain unchanged. Additional benefits include improved texture memory bandwidth and reduced fill rate pressure. Our technical approach uses texel density analysis and mipmap LOD bias adjustments for further gains. Contact us to discuss details.
Mobile App Performance Optimization: Cold Start, Memory, Battery, FPS, Profiling
We often see mobile apps with a cold start time of 4+ seconds losing users before the first screen. Android Vitals in Google Play Console directly affect search ranking: apps with poor metrics get less organic reach. Apple similarly monitors crash rate and launch time via MetricKit. Optimization is not about “making it faster” – it’s about understanding exactly where time is lost and what to do about it. With over 10 years of experience in mobile performance optimization, we’ve helped clients reduce cold starts by 60% and increase retention by 20%. Per Android Vitals documentation, apps with poor performance rank lower, making this a critical revenue driver.
How to Profile Mobile App Performance?
Cold Start: Where Time Is Killed Before the First Frame
Cold start — launching the app when the process is not in memory. On Android, this is the time from tapping the icon to Activity.onWindowFocusChanged(hasFocus = true). On iOS, from tap to viewDidAppear of the first screen.
Android: Main Thread Overloaded During Initialization
Application.onCreate() — the main enemy of fast start on Android. Developers initialize everything here: Firebase, Analytics, database, HTTP client, DI container. Each SDK adds 20–200 ms on the main thread.
Diagnostic tool: Android Studio Profiler → App Startup. Shows the initialization graph with time for each component. Alternative: Tracing.beginSection(“MyInitTag”) in code + systrace.
Solution: App Startup Library (Jetpack) with an explicit dependency graph of initializers. Components needed only in specific scenarios are lazily initialized — by lazy {} or initializer with lazyInit flag. Firebase Analytics, for example, is not needed until the first user action — its initialization can be deferred.
ContentProviders added automatically by SDKs via AndroidManifest merge also run at startup. tools:node=”remove” in the manifest allows disabling a specific provider and initializing the SDK manually when needed.
Another pitfall: Room.databaseBuilder().build() on the main thread. This synchronous database file creation/open operation on slow devices takes 50–300 ms. Move it to a coroutine with Dispatchers.IO, in ViewModel via viewModelScope.launch.
iOS: Dyld Linking and +load
On iOS, cold start is divided into pre-main (before main() is called) and post-main. Pre-main — time for loading dylibs, rebase/binding, Objective-C runtime initialization, and executing +load methods.
Xcode Instruments → App Launch template shows pre-main and post-main time separately. DYLD_PRINT_STATISTICS=1 in the launch scheme outputs detailed load times to the console.
Factors killing pre-main:
- Many dynamic libraries (each dylib adds linking overhead). CocoaPods adds a separate dylib per pod. Solution: Swift Package Manager with static linking (
type: .static) or use_frameworks! :linkage => :static in CocoaPods. Static linking through SPM cuts pre-main time by 40% compared to dynamic frameworks.
-
+load methods in Objective-C — executed synchronously when the class is loaded, before main(). Third-party SDKs may abuse this. +initialize — lazy alternative, called on first access to the class.
Post-main — application(_:didFinishLaunchingWithOptions:). Same story as on Android: synchronous initialization of everything. Use lazy var for services not needed immediately. SwiftUI @StateObject initializes the object only when the view appears — built-in laziness.
Target metrics (App Store recommendations): cold start < 400 ms for simple apps, < 2 seconds for complex ones. Warm start (process in memory, but Activity/Scene is recreated) — < 1 second. After optimization, we typically see cold start drop from 3.2s to 1.1s on mid-range devices.
Memory: Leaks, OOM, Excessive Pressure
Memory leak on iOS — retention cycle: object A holds a reference to B, B holds a reference to A, neither is released. Classic: Timer with self in closure without [weak self]. Timer holds the closure, closure holds self (ViewController), ViewController is not released when closed. Instruments → Leaks finds alive objects that should not be there.
On Android, garbage collector manages memory, but leaks still happen. Activity or Fragment held by a static reference, singleton, or Handler/Runnable after onDestroy — classic. LeakCanary is mandatory in debug builds. Add one dependency debugImplementation “com.squareup.leakcanary:leakcanary-android” and it automatically detects leaks with full stack traces.
OutOfMemoryError is most often due to image loading. Bitmap in memory occupies width × height × 4 bytes. An image 4000×3000 px — 48 MB in memory, regardless of file size on disk. Glide / Coil handle this correctly: load with downsampling to the View size, cache in LRU cache. Loading into ImageView without Glide/Coil via BitmapFactory.decodeFile is a path to OOM on devices with 2 GB RAM. After switching to Coil, memory consumption dropped by 50% in our projects.
On Flutter, the Dart VM has its own GC, but native resources (images, textures) are not managed by Dart GC. Image.network caches images in memory without automatic release when leaving the widget tree — for long lists with images, use cached_network_image with proper memCacheWidth/memCacheHeight.
Why Does Cold Start Take So Long? Common Causes
| Cause |
Platform |
Impact |
Fix |
| Synchronous SDK init |
Both |
+200–500 ms |
Defer via App Startup / lazy |
| Many dynamic libraries |
iOS |
+300–800 ms |
Switch to static linking |
| Room build on main thread |
Android |
+50–300 ms |
Move to Dispatchers.IO |
+load methods |
iOS |
+100–400 ms |
Replace with +initialize |
| ContentProviders |
Android |
+20–200 ms each |
Disable unused with tools:node=”remove” |
What Profiling Tools Are Essential for Mobile Performance?
FPS and UI Performance
60 FPS — 16.67 ms per frame. 120 FPS (ProMotion) — 8.33 ms. Anything taking longer on the main thread causes jank.
Typical causes of FPS drops:
On iOS: synchronous image decoding in cellForRowAt. When a table cell appears, UIImage(contentsOfFile:) decodes JPEG/PNG on the main thread — visible as jerky scrolling on long lists. Solution: UIImage.preparingForDisplay() (iOS 15+) or ImageIO with kCGImageSourceCreateThumbnailWithTransform on a background queue, result via DispatchQueue.main.async.
On Android: RecyclerView.Adapter.onBindViewHolder with synchronous operations. Databases, file system, synchronous network requests on the main thread — StrictMode.ThreadPolicy with detectAll().penaltyLog() in debug builds will show all violations.
On Flutter: build() method is called frequently; it must be cheap. setState() on a top-level widget rebuilds the entire tree. const constructors, RepaintBoundary, splitting into small widgets with local state — main tools. Flutter DevTools → Performance shows janky frames (red) with causes.
Compose profiling: Recomposition Highlighter and tracing via Trace.beginSection in @Composable. Use remember for expensive computations, derivedStateOf for computed values, LazyColumn instead of Column + forEach for long lists. Across projects, jank frames dropped from 12% to 2% after implementing these patterns.
Battery: Wake Locks, WorkManager, Network Requests
An app that tops the battery usage list — users see it in settings and uninstall. Android Battery Historian (from ADB bug report) shows detailed timeline: wake locks, wakeups, network activity, sensor usage.
Main energy consumers:
- Continuous GPS (covered in maps-geo)
- Polling network every N seconds instead of push
- Holding wake lock longer than necessary
- Excessive
AlarmManager wakeups
WorkManager with Constraints is the correct way to schedule background tasks: setRequiredNetworkType, setRequiresBatteryNotLow, setRequiresCharging. The OS batches tasks and executes them at convenient times.
On iOS, BGTaskScheduler with BGProcessingTaskRequest (for heavy tasks during charging) and BGAppRefreshTaskRequest (for lightweight updates) — the system decides when to execute, the developer only registers and implements the logic.
Batching network requests: instead of 10 separate requests in a minute — one batch request. Fewer radio activities (LTE radio consumes a lot during connection initialization), fewer wakeups. This typically cuts battery usage by 30% in network-heavy apps.
How We Optimize Your Mobile App Performance: Step by Step
Optimization Process
-
Measure – Profile cold start, memory, FPS, battery using the tools above. Obtain baseline numbers (e.g., cold start 3.2s, memory footprint 180 MB, 12% jank frames).
-
Analyze – Identify top 3 bottlenecks by impact. For a typical e‑commerce app, image loading and SDK init are priority.
-
Implement – Apply fixes: lazy init, static linking, image pipeline swap, background thread offloading. We deliver code changes with diff reports.
-
Test – Profile again; compare before/after numbers. Validate on real devices (including low-end).
-
Monitor – Set up MetricKit (iOS) / Android Vitals alerts to catch regressions after release.
Deliverables:
- Detailed profiling report with before/after metrics
- Annotated code diffs for each optimization
- Configuration recommendations (e.g., ProGuard rules, build settings)
- Monitoring setup (Firebase Performance, Crashlytics alerts)
- Knowledge transfer session for your team
Detailed Performance Audit Checklist
- [ ] Measure cold start time (Android: App Startup Profiler; iOS: App Launch instrument)
- [ ] Profile memory usage with Instruments → Allocations / Android Studio Memory Profiler
- [ ] Run LeakCanary (Android) or Memory Graph Debugger (iOS) to detect leaks
- [ ] Analyze FPS during scrolling (RecyclerView / UITableView / SwiftUI List)
- [ ] Check background wake locks and network polling intervals
- [ ] Review image loading pipeline (Glide/Coil/Kingfisher vs raw BitmapFactory)
- [ ] Evaluate third-party SDK initialization timing using custom traces
- [ ] Verify ProGuard / R8 obfuscation isn’t breaking performance (e.g., reflection)
- [ ] Test on a representative low-end device (e.g., Samsung Galaxy A21, iPhone SE)
Estimated Timeline
| Scope |
Duration |
| Performance audit (existing app) |
3–5 working days |
| Optimizations (tier 1 – low‑hanging fruit) |
1–2 weeks |
| Full optimization campaign (including architecture changes) |
2–8 weeks |
Costs are calculated individually based on app complexity and current codebase state. Contact us for a project estimate and performance review.
Profiling Tools Reference
| Platform |
Tool |
What It Shows |
| iOS |
Xcode Instruments (Time Profiler) |
CPU, call stack, hot methods |
| iOS |
Allocations |
Live objects, memory peaks |
| iOS |
Leaks |
Retention cycles |
| iOS |
MetricKit |
Production metrics (crash rate, hang rate, launch time) |
| Android |
Android Profiler |
CPU, Memory, Network, Energy |
| Android |
Systrace / Perfetto |
System-level traces |
| Android |
LeakCanary |
Memory leaks |
| Android |
Battery Historian |
Energy consumption |
| Flutter |
Flutter DevTools |
Recomposition, frame rendering, memory |
| Flutter |
Dart Observatory |
Dart VM profiling |
MetricKit on iOS is especially valuable: real data from user devices, not simulator. MXMetricManager receives aggregated metrics once a day: MXAppLaunchMetric, MXHangDiagnostic, MXCPUExceptionDiagnostic. Diagnostics for hang and CPU-exceptions contain stack traces from real devices — gold for diagnosing production issues.
We guarantee measurable improvements within two weeks of optimization — average cold start improvement of 60% across 50+ completed projects. Get in touch for a tailored performance review.