IaaS vs PaaS vs SaaS and Cloud Cost Optimization for Data Teams

Data Engineering for Beginners by Chisom Nwokwu (ISBN 9781394325412, Wiley 2025)

Previous: Cloud Data Engineering Basics

The second half of Chapter 12 answers two questions every data engineer eventually faces: how much of the stack do you want to manage yourself, and how do you keep the cloud bill from giving finance a heart attack?


IaaS, PaaS, and SaaS

Nwokwu uses a pyramid diagram to show the three service models. The higher you go, the less you manage.

Infrastructure as a Service (IaaS)

You get virtual machines, storage, and networking. You install the OS, configure Spark, manage Airflow upgrades, handle security patches. Full control, full responsibility.

Example: spinning up EC2 instances, installing Python and PySpark, writing custom ETL scripts. Good when you need a tailored setup and have the team to maintain it.

Platform as a Service (PaaS)

The provider handles infrastructure and runtime. You focus on building pipelines and writing transformations.

Examples: Google Cloud Composer (managed Airflow), Azure Data Factory, AWS Glue Studio. Less server babysitting, faster deployment. Trade-off: less control over the underlying environment.

Software as a Service (SaaS)

Ready-to-use applications over the internet. You configure workflows through a UI. The provider handles everything else.

Examples: Snowflake, BigQuery, Fivetran. Fastest to get running. Least flexible. Your data lives on their servers, which raises privacy questions in regulated industries.

How to choose

Nwokwu walks through seven decision factors:

  1. Organization needs - custom platform? IaaS. Fast pipeline development? PaaS. Quick dashboards? SaaS
  2. Cost - IaaS looks cheap until you count engineering time; SaaS has predictable subscriptions but per-query costs can spike
  3. Security - IaaS gives full control but you own the work; SaaS offloads compliance but you trust the vendor
  4. Scalability - IaaS and PaaS scale well; SaaS scales users easily but limits performance tuning
  5. Time to deploy - SaaS is fastest, IaaS is slowest
  6. Integration - IaaS fits legacy systems best; SaaS is easy when connectors exist
  7. Long-term strategy - start with SaaS for quick wins, migrate to PaaS or IaaS when you need more control

Her bottom line: most real organizations use a hybrid of all three. IaaS for custom infrastructure, PaaS for orchestration, SaaS for analytics and BI.

Serverless, managed, and self-managed

Beyond service models, Nwokwu introduces three management models:

Serverless - write functions, cloud runs them on events. No servers to provision. Perfect for file-upload triggers, lightweight validations, and alerts. Downsides: cold start latency, execution time limits, harder debugging.

Managed - provider handles infrastructure, scaling, and failover. BigQuery, Snowflake, managed Airflow. Less ops work, but vendor lock-in is real.

Self-managed - you run Spark on VMs, patch Hadoop yourself, tune everything. Maximum flexibility, maximum maintenance.

She recommends mixing all three in one architecture. Serverless for event-driven steps, managed for warehouses, self-managed for custom performance-sensitive workloads.

Cost optimization (the part finance cares about)

This section felt like advice from someone who has seen a surprise AWS bill.

Pricing models

  • On-demand - flexible, pay per second. Good for dev and testing. Expensive for 24/7 production jobs.
  • Reserved instances - commit for 1-3 years, get a big discount. Use for steady ETL pipelines.
  • Spot instances - cheap leftover capacity, but AWS can reclaim it anytime. Great for fault-tolerant batch jobs with checkpointing.

Practical tips

  • Rightsize - do not give a 2-vCPU job an 8-vCPU machine. Monitor actual usage.
  • Schedule smart - run batch jobs off-peak when spot pricing drops.
  • Storage tiers - hot storage for active data, cold/archive for old logs. Set lifecycle policies to auto-move data after 30 or 90 days.
  • Shut down idle resources - dev clusters running overnight are wasted money.
  • Serverless for light tasks - pay only for execution time, not idle servers.
  • Monitor and alert - set budgets before finance comes knocking.

Nwokwu’s line that stuck: “Treat cloud infrastructure like a utility. Turn it off when you are not using it.”

My take

The service model section could feel abstract, but Nwokwu grounds every choice in data engineering scenarios. “You need custom Spark tuning” maps to IaaS. “You want to ship a pipeline this week” maps to PaaS. “You need a warehouse yesterday” maps to SaaS.

The cost section is the most immediately useful part of the chapter. Rightsizing and lifecycle policies are boring topics that save real money. If you are interviewing for a data engineering role, being able to talk about spot instances for batch jobs and storage tiering shows you think beyond just making pipelines work.


Next: Building a Data Engineering Career