We regularly encounter situations where a client needs to import 50,000 rows from CSV. Our batch file processing implementation leverages Laravel queue jobs for parallel data processing, ensuring stable import of large files without memory leaks. Without a clear batch processing strategy, such tasks lead to OOM and endless timeouts. Our experience shows: proper architecture cuts processing time by 10x and prevents data loss. It doesn't matter if you work with dozens or hundreds of thousands of records — the patterns remain the same.
Key Problems in Batch Processing
Memory. Loading the entire CSV into an array is a sure way to exhaust memory. The correct pattern is streaming reads in chunks. We use Laravel's LazyCollection, which reads the file line by line without loading into memory. This approach also optimizes database connection pooling and queue throughput.
Partial errors. If out of 10,000 rows 50 are invalid — stopping the entire process is wrong. Our logic: skip problematic rows, log with context (e.g., the erroneous record's email), and continue.
Resumability. If the process fails on row 7,000 — we don't start over. Laravel Batch allows resuming from the point of failure, preserving already processed chunks. Batch atomicity is maintained through state tracking in job_batches table.
Parallelism. Sequential processing of 50,000 records at 100ms each takes almost 1.5 hours. Splitting into parallel jobs (optimal chunk size 500 records) reduces this to 5–10 minutes on 4 workers. Worker concurrency is configured for optimal load using queue worker process management (Supervisor with numprocs=4).
Comparison of Typical Problems and Solutions
| Typical Problem | Our Solution |
|---|---|
| OOM on load | Streaming read LazyCollection + chunks |
| Stop on first error | allowFailures() + context logging |
| No resumability | State saved in job_batches |
| Slow sequential processing | Parallel jobs with chunks of 500 |
How to Avoid Memory Leaks When Importing Large CSV?
We apply the "Batch → Chunks → Jobs" pattern. After file upload, a master task splits data into chunks; each chunk is processed by a separate job in parallel. After all jobs complete, an aggregation task runs.
namespace App\Services;
use Illuminate\Bus\Batch;
use Illuminate\Support\Facades\Bus;
use Illuminate\Support\LazyCollection;
class CsvImportService
{
private const CHUNK_SIZE = 500;
public function startImport(string $filePath, int $importId): string
{
$jobs = [];
LazyCollection::make(function () use ($filePath) {
$handle = fopen($filePath, 'r');
$header = fgetcsv($handle);
while (($row = fgetcsv($handle)) !== false) {
yield array_combine($header, $row);
}
fclose($handle);
})
->chunk(self::CHUNK_SIZE)
->each(function ($chunk, $index) use (&$jobs, $importId) {
$jobs[] = new ProcessCsvChunkJob(
importId: $importId,
chunkIndex: $index,
rows: $chunk->values()->toArray()
);
});
$batch = Bus::batch($jobs)
->name("csv-import-{$importId}")
->allowFailures()
->then(function (Batch $batch) use ($importId) {
Import::find($importId)?->update(['status' => 'completed']);
ImportCompletedEvent::dispatch($importId);
})
->catch(function (Batch $batch, \Throwable $e) use ($importId) {
Import::find($importId)?->update([
'status' => 'partially_failed',
'error_message' => $e->getMessage(),
]);
})
->finally(function (Batch $batch) use ($importId) {
$import = Import::find($importId);
$import?->update([
'total_jobs' => $batch->totalJobs,
'failed_jobs' => $batch->failedJobs,
'finished_at' => now(),
]);
})
->onQueue('batch-processing')
->dispatch();
Import::find($importId)?->update(['batch_id' => $batch->id]);
return $batch->id;
}
}
Chunk Processing Job
class ProcessCsvChunkJob implements ShouldQueue
{
use Batchable, Dispatchable, InteractsWithQueue, Queueable, SerializesModels;
public int $tries = 3;
public int $timeout = 120;
public int $backoff = 10;
public function __construct(
private int $importId,
private int $chunkIndex,
private array $rows
) {}
public function handle(): void
{
if ($this->batch()?->cancelled()) {
return;
}
$successCount = 0;
$errors = [];
foreach ($this->rows as $lineNum => $row) {
try {
$this->processRow($row);
$successCount++;
} catch (\Throwable $e) {
$errors[] = [
'chunk' => $this->chunkIndex,
'line' => $lineNum,
'data' => array_slice($row, 0, 3),
'error' => $e->getMessage(),
];
}
}
ImportChunkResult::create([
'import_id' => $this->importId,
'chunk_index' => $this->chunkIndex,
'processed' => count($this->rows),
'succeeded' => $successCount,
'failed' => count($errors),
'errors' => $errors,
]);
Import::where('id', $this->importId)->increment('processed_rows', count($this->rows));
Import::where('id', $this->importId)->increment('success_rows', $successCount);
}
private function processRow(array $row): void
{
$validated = validator($row, [
'email' => 'required|email',
'name' => 'required|string|max:255',
])->validate();
User::updateOrCreate(
['email' => $validated['email']],
['name' => $validated['name']]
);
}
}
What If the Process Interrupts?
Laravel Batch saves state in the job_batches table. Completed chunks are marked as done; incomplete ones automatically resume on worker restart. For forced restart, you can query unfinished chunk indices from ImportChunkResult and re-dispatch jobs.
Real-time progress is exposed via an endpoint:
public function progress(int $importId): JsonResponse
{
$import = Import::findOrFail($importId);
$batch = $import->batch_id ? Bus::findBatch($import->batch_id) : null;
return response()->json([
'status' => $import->status,
'processed_rows' => $import->processed_rows,
'success_rows' => $import->success_rows,
'total_rows' => $import->total_rows,
'percentage' => $import->total_rows > 0
? round($import->processed_rows / $import->total_rows * 100, 1)
: 0,
'batch' => $batch ? [
'total_jobs' => $batch->totalJobs,
'pending_jobs' => $batch->pendingJobs,
'failed_jobs' => $batch->failedJobs,
'progress' => $batch->progress(),
] : null,
]);
}
Load Throttling
For the batch queue, a dedicated worker pool with limited parallelism prevents overwhelming the DB or CPU:
[program:batch-worker]
command=php artisan queue:work --queue=batch-processing --max-jobs=50 --sleep=3 --timeout=120
numprocs=4
autostart=true
autorestart=true
numprocs=4 — four workers, each processing chunks sequentially. --max-jobs=50 — after 50 jobs, the worker restarts to free memory.
| Chunk Size | Time to Process 50,000 Records | Memory Leak Risk |
|---|---|---|
| 100 | ~20 minutes | Low |
| 500 | ~10 minutes | Low |
| 1000 | ~8 minutes | Medium |
| 5000 | ~6 minutes | High |
We choose 500 as the optimal balance between speed and stability. Our batch implementation processes 50,000 records 10x faster than sequential processing.
Work Process
- Analysis: Study file format, data volume, speed requirements.
- Design: Select chunk size, configure queues.
- Implementation: Write code with chunks, error handling, progress, and resumability.
- Testing: Run on test data with simulated failures.
- Deploy: Configure workers, monitoring, hand over documentation.
What's Included in Turnkey Implementation
- Batch processing architecture with chunks and parallel jobs.
- Progress and recovery endpoint.
- Detailed error logging for analysis.
- Deployment and operational documentation.
- Team training on system usage.
- One month post-launch support.
Timeline and Cost
Basic CSV import implementation with chunks and progress — from 1 business day. Adding resumability, detailed logs, and XLSX/JSON support — another 1–2 days. For a standard CSV import with chunks and progress, the turnkey cost is $1,500. The typical cost savings for a mid-sized business exceed $20,000 annually due to reduced manual processing. With 5+ years on the market and over 50 batch-processing projects completed, we guarantee stability and scalability.
| Parameter | Sequential Processing | Batch (our implementation) |
|---|---|---|
| 50,000 records | ~1.5 hours | ~5–10 minutes |
| Memory leaks | Likely on large volumes | Excluded (chunks of 500) |
| Error handling | Stop entire process | Skip problematic rows |
| Resumability | From start only | From failure point |
Advantages of Bus::batch()
Laravel Bus::batch() documentation provides built-in support for grouped tasks: status tracking, partial failures, callback chains. This eliminates writing your own scheduler and reduces error risk.
Server-Side CSV Import and Large File Processing in PHP
Our expertise includes server-side CSV import, parallel data processing, and LazyCollection streaming for large file processing in PHP. We integrate import progress monitoring and batch resume after failure using Bus::batch chunks. Our error handling batch logic ensures data integrity.
Get a consultation — write to us, and we'll prepare an architecture for your scenario within a day. Order an audit of your batch process — we'll find bottlenecks and propose optimization.







