Serverless File Processing: Lambda & S3 Trigger Turnkey

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Serverless File Processing: Lambda & S3 Trigger Turnkey
Medium
~2-3 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    947

Serverless File Processing: Lambda & S3 Trigger Turnkey

When uploading images to S3 you need automatic thumbnail generation, but Lambda struggles with large files. You need a fault-tolerant architecture with queues and monitoring. Our team has over 10 years of cloud experience and has delivered 200+ serverless projects, ensuring robust architecture. In this article we break down a typical architecture, error handling, and pitfalls.

Lambda + S3 trigger is a classic serverless pattern for file processing. A file is uploaded to S3, the event automatically triggers Lambda, which processes the file and saves the result. No permanently running server needed, scaling is automatic. Without a proper queue scheme, however, you risk losing events. Cost savings are significant: for example, processing 1 million images per month costs around $50 with Lambda, compared to $200 with a fixed server. Let's explore how to build production-grade processing with queues, DLQ, and monitoring.

Common Tasks and Basic Architecture

Typical tasks: generating thumbnails on image upload, converting video to different formats and resolutions, processing CSV/Excel files with data import to a database, PDF generation from templates, antivirus scanning of uploaded files, OCR and text extraction from documents, data transformation (XML → JSON, normalization).

File Type Examples Recommended Approach
Images JPG, PNG, WebP Lambda + Pillow
Documents PDF, DOCX Lambda + pdfminer, python-docx
Video MP4, AVI AWS MediaConvert
Spreadsheets CSV, XLSX Lambda + pandas (streaming)
Archives ZIP, RAR Lambda + zipfile (streaming)

Basic architecture:

[User] → S3 upload → [S3 Event Notification]
                                ↓
                        [Lambda Function]
                                ↓
                [Processed file → S3 Output]
                [Metadata → DynamoDB]
                [Notification → SQS/SNS]
# S3 bucket for incoming files
resource "aws_s3_bucket" "uploads" {
  bucket = "myapp-uploads"
}

# S3 bucket for processed files
resource "aws_s3_bucket" "processed" {
  bucket = "myapp-processed"
}

# Combined configuration: S3 → SNS → SQS
resource "aws_s3_bucket_notification" "upload_trigger" {
  bucket = aws_s3_bucket.uploads.id

  topic {
    topic_arn = aws_sns_topic.file_events.arn
    events    = ["s3:ObjectCreated:*"]
  }
}

resource "aws_sns_topic_subscription" "to_sqs" {
  topic_arn = aws_sns_topic.file_events.arn
  protocol  = "sqs"
  endpoint  = aws_sqs_queue.file_processing.arn
}

Scaling Processing Under Peak Load

Lambda scales automatically, but a sudden spike in uploads can cause throttling. We use an SQS queue for buffering: S3 → SNS → SQS → Lambda. This guarantees no event gets lost. When the concurrent execution limit is exceeded, the queue accumulates messages, and Lambda processes them as capacity frees up. For critical tasks we configure Reserved Concurrency. Compared to a fixed server cluster, Lambda is 5x faster to deploy and scales instantly.

Lambda Handler and Large Files

import boto3
import json
import os
from urllib.parse import unquote_plus
from PIL import Image
import io

s3 = boto3.client('s3')

def handler(event, context):
    results = []
    
    for record in event['Records']:
        bucket = record['s3']['bucket']['name']
        key = unquote_plus(record['s3']['object']['key'])
        
        try:
            result = process_image(bucket, key)
            results.append({'key': key, 'status': 'success', **result})
        except Exception as e:
            print(f"Error processing {key}: {e}")
            results.append({'key': key, 'status': 'error', 'error': str(e)})
    
    return results

def process_image(bucket: str, key: str) -> dict:
    # Download original
    obj = s3.get_object(Bucket=bucket, Key=key)
    image_data = obj['Body'].read()
    
    image = Image.open(io.BytesIO(image_data))
    
    thumbnails = {}
    for size_name, (width, height) in [('sm', (150, 150)), ('md', (400, 400)), ('lg', (800, 800))]:
        thumb = image.copy()
        thumb.thumbnail((width, height), Image.LANCZOS)
        
        buffer = io.BytesIO()
        thumb.save(buffer, format=image.format or 'JPEG', quality=85)
        buffer.seek(0)
        
        output_key = key.replace('images/', f'thumbnails/{size_name}/')
        s3.put_object(
            Bucket=os.environ['OUTPUT_BUCKET'],
            Key=output_key,
            Body=buffer,
            ContentType=f'image/{(image.format or "JPEG").lower()}'
        )
        thumbnails[size_name] = output_key
    
    return {'thumbnails': thumbnails, 'original_size': image.size}

Lambda has limits: /tmp up to 10 GB, timeout up to 15 minutes, memory up to 10 GB. For files >100 MB we use streaming, processing data in chunks without loading it all into memory.

Why SNS+SQS Is More Reliable Than Direct Lambda Invocation?

S3 event notifications do not support DLQ directly. If Lambda errors, the event is lost. The SNS+SQS scheme guarantees delivery: failed invocations go to DLQ after N retries. You can analyze and reprocess them. This follows AWS best practices for production. According to AWS documentation, this pattern ensures 99.9% event delivery reliability.

More about DLQ configuration

For the SQS queue, a redrive policy is set: after, for example, 5 failed attempts the message moves to DLQ. This ensures events are not lost and errors can be analyzed.

Lambda vs ECS: What to Choose?

Criterion Lambda ECS Fargate
Max execution time 15 min Unlimited
Max memory 10 GB 30 GB
Cost Per millisecond Per vCPU/hour
Scaling Instant 30–60 sec
Optimal for Short tasks (<15 min) Long tasks, high memory

For image and document processing, Lambda is simpler and cheaper. For video transcoding we use AWS MediaConvert.

Lambda Function Configuration for File Processing

resource "aws_lambda_function" "processor" {
  filename         = "processor.zip"
  function_name    = "file-processor"
  role             = aws_iam_role.processor.arn
  handler          = "handler.handler"
  runtime          = "python3.12"
  timeout          = 300
  memory_size      = 1024
  
  ephemeral_storage {
    size = 2048
  }
  
  environment {
    variables = {
      OUTPUT_BUCKET = aws_s3_bucket.processed.bucket
    }
  }
}

Monitoring and Metrics

  • Number of files processed per hour (typical throughput: 1000 files/min)
  • Average processing time by file type (e.g., image resizing averages 0.3 sec)
  • Error rate + DLQ contents (target: <0.1% error rate)
  • Lambda duration distribution

CloudWatch Dashboard with these metrics plus an alert on DLQ growth.

What's Included and Estimated Timelines

  • Development of Lambda function in Python (or Node.js/Go) for your file type
  • Infrastructure in Terraform / Pulumi with S3, SNS, SQS, DLQ
  • CloudWatch monitoring setup + alerts
  • Code review and load testing up to 1000 files/min
  • Architecture documentation and deployment instructions
  • Client team training (1-2 hours)
  • 2 weeks post-delivery support

Estimated timelines:

  • Requirements analysis and prototype — 2-3 days
  • Full pipeline with DLQ, monitoring — 4-6 days
  • Integration with your application (API Gateway, Cognito — optional) — 2-5 days

Actual timelines depend on processing complexity. Contact us for a one-day project assessment. Request a consultation and get a free preliminary estimate with no obligation.

Why Serverless Development? The Real Economics and Technical Trade-offs

Serverless does not mean "without servers". Servers exist—you just don't manage them. It's more accurate to think of it as "without server management": no OS patching, no nginx configuration, no disk space monitoring. The function receives an event, processes it, and returns a response. The provider decides where to run it. Мы занимаемся serverless-архитектурой более 5 лет и реализовали 30+ проектов на AWS Lambda, Vercel Functions и Cloudflare Workers. Гарантируем, что ваша система масштабируется без переплат — при условии правильного выбора платформы и оптимизации холодного старта.

Platform Cold Start (Node.js) State Management Bundle Size Limit Best For
AWS Lambda 200ms–1.5s (VPC: до 10s) External (DynamoDB, S3) 250MB (with layers) Complex event‑driven, enterprise
Vercel Functions ~300ms (50ms with Edge) Edge Config, KV 4MB (Edge), 50MB (Serverless) Next.js, JAMstack, middleware
Cloudflare Workers <1ms Durable Objects, KV, D1 1MB (worker code) Global low‑latency, real‑time

Cold start — Lambda's main pain point on Node.js. In VPC, cold start reached 10 seconds before recent improvements. For production functions with latency requirements: Provisioned Concurrency (keeps instances warm), SnapStart for Java, minimize bundle via tree-shaking. Our typical optimization reduces cold start from 3.2s to 400ms.

Practical case: an image processing function (resize, WebP conversion, upload to S3). Bundle with sharp was 40MB due to native binaries. Solution: Lambda Layer with sharp, main function 800KB. Cold start dropped from 3.2s to 400ms. Lambda Layers — shared dependencies between functions. Up to 5 layers per function, each up to 250MB. Standard practice: layer with heavy dependencies (sharp, puppeteer, ffmpeg), layer with common business logic. Infrastructure for Lambda via AWS CDK or Terraform. SAM — for beginners, CDK — for serious projects with type safety.

Edge Runtime is fundamentally different: the function runs on a V8 isolate in the nearest Vercel CDN point (120+ regions). No cold start as such — the isolate starts in ~0ms. But strict limitations: no Node.js API (fs, crypto via Web API), no database access via TCP (only via HTTP API), bundle size up to 4MB. Edge Runtime is ideal for: middleware (auth check, redirect, A/B test), response transformations, geolocation logic, Edge Config. Not suitable for: accessing PostgreSQL, heavy computations, file system operations.

Cloudflare Workers run on V8 isolates in 300+ points of presence. Latency for the user is literally the nearest data center. Cold start < 1ms. Workers Durable Objects solve the state problem at the edge: each Durable Object is a single coordination point, running in one region. Ideal for: game rooms, real-time documents, rate limiting without races. Workers KV — eventually consistent storage. Writes propagate to all regions in ~60 seconds. Not suitable for financial transactions, suitable for configs, feature flags, cache. D1 — SQLite on the edge. Works great on a single read replica, write latency depends on distance to primary region. Not ideal for global write-heavy applications.

Ecosystem: Hono.js — a minimalist router that works on Workers, Deno, Bun, Node.js. Good choice if you need unified code for edge and server.

Vendor lock-in — a real problem. Lambda-specific code (handler signature, Lambda context) is hard to port. Hono.js, Remix, or adapters like @hono/node-server help keep logic portable. Мы проектируем абстракции, позволяющие сменить провайдера с минимальными изменениями.

How We Optimize Cold Start in AWS Lambda?

Cold start is Lambda's worst enemy. Here’s a step‑by‑step optimisation checklist we apply:

  1. Minimise bundle size — tree‑shake dependencies, use Lambda Layers for native binaries (sharp, puppeteer). Target < 1MB.
  2. Enable Provisioned Concurrency for latency‑critical functions — costs extra but cuts cold start to near zero.
  3. Use SnapStart for Java (Lambda) — reduces init time by 90%+.
  4. Avoid VPC unless necessary — if you need VPC, use AWS PrivateLink or Elastic Network Interface optimisation.
  5. Warm‑up strategies — scheduler pinging function every 5 minutes (but only for low‑volume functions, otherwise Provisioned Concurrency cheaper).

Result: our clients typically see cold start drop from 2–4s to under 500ms. For a fintech API handling 50k requests/day, that means 3 fewer seconds of latency per request during peak scale.

When Does Serverless Not Fit? Cost Comparison

Serverless saves money when traffic is unpredictable or sparse — up to 70% reduction compared to dedicated servers. But it becomes expensive under constant high load. Example: a function processing 1 million requests/day at 300ms each costs about $100–200/month on Lambda. Equivalent EC2 instance might cost $50/month. For such steady workloads, Fargate or EC2 is cheaper.

Long computations (>15 min on Lambda, >30s on Vercel) require Fargate or a regular server. WebSocket server with state — no persistent process. Tasks with frequent disk access — ephemeral storage, /tmp on Lambda 512MB–10GB.

What’s Included in Serverless Development Service?

Мы предлагаем serverless-разработку под ключ. В каждый проект входит:

  • Архитектурная документация (схема event‑driven потоков, выбор платформы, justification).
  • Реализация функций с unit‑ и integration‑тестами.
  • CI/CD pipeline (GitHub Actions / GitLab CI) с preview‑деплоями.
  • Infrastructure as Code (Terraform / AWS CDK / Pulumi).
  • Мониторинг и observability (OpenTelemetry, structured logging, distributed tracing).
  • 30‑дневная пост‑релизная поддержка и оптимизация производительности.

Typical Mistakes in Serverless Development and How We Avoid Them

  • Ignoring cold start — we measure and budget for it from day one.
  • Over‑engineering state — many teams try to use Workers Durable Objects for simple caching; KV is often enough.
  • No distributed tracing — without trace IDs across SQS › Lambda › DynamoDB streams, debugging is blind. We integrate AWS X‑Ray or OpenTelemetry automatically.
  • Underestimating cost at scale — we simulate load patterns and compare serverless vs. container costs before committing.

Закажите serverless архитектуру под ключ — свяжитесь с нами для бесплатной оценки вашего проекта. Сроки: от 2 недель для MVP, до 10 недель для миграции монолита. Стоимость рассчитывается индивидуально, ориентировочно от $2,000 до $15,000 в зависимости от сложности.