What Is Serverless Computing and How Does It Work?

Serverless computing is a cloud execution model where the cloud provider dynamically manages all the infrastructure, from provisioning servers to scaling capacity and handling maintenance. You write functions, deploy them, and the platform runs them only when triggered by an event, charging you solely for the compute time consumed. A widely cited Berkeley paper describes it as an evolution that “parallels the transition from assembly language to high-level programming languages,” because it strips away virtually all the system administration work traditionally needed to use the cloud.1arXiv. Cloud Programming Simplified: A Berkeley View on Serverless Computing The name is somewhat misleading, though, and the trade-offs are real enough that understanding them matters before you commit.

What “Serverless” Actually Means

Servers still exist. You just never see, configure, or patch them. The term describes the developer’s experience, not the physical reality. You write a piece of application logic, often called a function, and upload it to a platform like AWS Lambda, Azure Functions, or Google Cloud Functions. That function sits dormant until something triggers it: an HTTP request, a database change, a file upload, a message arriving in a queue. The platform spins up an environment, runs your code, returns the result, and tears everything down. If nobody calls the function, nothing runs and you pay nothing.

The core principles are straightforward. There is no server management on your end. Scaling is automatic: if ten thousand requests arrive simultaneously, the platform launches enough instances to handle them. Billing is granular, typically measured in milliseconds of execution time and the memory allocated during that window. And because the platform manages redundancy and failover, you inherit a degree of fault tolerance without engineering it yourself.2Journal of Emerging Trends in Computer Science and Applications. Serverless Computing: Event-Driven Architectures for Agile Development

This pay-per-use model is what separates serverless from renting a virtual machine. A traditional cloud server costs money whether it is handling requests or sitting idle at 3 a.m. A serverless function costs zero when idle. For workloads that are bursty or unpredictable, like a webhook processor or an image-resizing pipeline, that difference in billing can be dramatic.

The Cold Start Problem

The single most talked-about drawback of serverless is cold start latency. When a function has not been called recently, the platform has to set up a fresh execution environment: loading the runtime, pulling your code, initializing dependencies. That setup time can add hundreds of milliseconds, sometimes more, before your function actually begins its work. For a background data-processing job, nobody notices. For an API endpoint serving a user-facing app, those extra milliseconds feel sluggish.

Researchers have spent considerable effort trying to solve this. A systematic review of the cold start literature classifies the approaches into several categories, including caching strategies, application-level optimizations, and machine-learning-based predictions that try to anticipate when a function will be called and pre-warm the environment before the request arrives.3ACM Computing Surveys. Cold Start Latency in Serverless Computing: A Systematic Review, Taxonomy, and Future Directions Despite all of this work, the review concludes that cold starts remain an unresolved research area. The fundamental tension is baked into the model: zero scaling (where you pay nothing when idle) means the environment must be torn down at some point, and rebuilding it takes time.

In practice, teams handle this in a few ways. Some cloud providers offer “provisioned concurrency,” which keeps a set number of environments warm at all times, though this sacrifices part of the cost advantage. Others write their functions in languages with lighter startup footprints. And some workloads simply tolerate the occasional slow response, because the cost savings and operational simplicity outweigh the latency penalty.

Statelessness and Why It Complicates Things

Serverless functions are stateless by design. Each invocation starts from a blank slate. It has no memory of the previous request, no local file that persists between calls, no session stored in RAM. This is a feature when it comes to scaling, because any instance can handle any request without worrying about what came before. But it is a real headache when your application needs to track state, run multi-step workflows, or maintain data consistency across several functions.

Suppose you have an e-commerce checkout that involves reserving inventory, charging a payment method, and sending a confirmation email. Each step is a separate function. If the payment succeeds but the email function fails, you need a way to handle that partial failure gracefully. In a traditional server, you might wrap all three in a database transaction. In serverless, the functions are independent and potentially running on different machines.

Researchers have proposed programming models that bring transactional guarantees to serverless workflows. One approach implements distributed transactions with two-phase commit for strict consistency, and Saga patterns for cases where eventual consistency is acceptable. These models leverage stateful dataflow engines to provide exactly-once processing guarantees underneath the stateless function layer.4Information Systems. Transactions across serverless functions leveraging stateful dataflows The tooling exists, but it adds complexity that you would not need in a more traditional architecture. If your application is inherently stateful and transactional, serverless may not be the most natural fit, or at least not without significant orchestration work.

A New Kind of Attack: Denial of Wallet

Pay-per-use billing creates a security vulnerability that did not exist in the old world of fixed-price servers. If an attacker floods your serverless functions with fake requests, the platform dutifully scales up to handle them all and sends you the bill. This is called a Denial of Wallet (DoW) attack, and it is essentially forced financial exhaustion rather than forced downtime.5Journal of Information Security and Applications. Denial of wallet—Defining a looming threat to serverless computing

Traditional DDoS attacks try to overwhelm a service so it stops responding. A Denial of Wallet attack does not need to take anything offline. It just needs to push your usage past whatever budget threshold hurts. An organization could wake up to a cloud bill that exceeds its contracted service quotas, sometimes by a staggering amount.6International Journal of Information Security. Entropy-based detection of denial of wallet attacks in serverless architectures

Defending against this requires a different mindset than traditional security. Rate limiting, budget alerts, spending caps, and anomaly detection on invocation patterns all help. Some teams set hard spending limits so the platform stops serving traffic beyond a threshold, though that trades financial risk for availability risk. Researchers have proposed entropy-based detection methods that analyze traffic patterns to distinguish legitimate bursts from artificially generated ones, but this remains an active area of development. The broader point is that serverless shifts certain risks from the operational domain into the financial domain, and your security strategy needs to follow.

Autoscaling Is Not as Simple as “Infinite Scale”

Marketing materials for serverless platforms emphasize automatic scaling, and to be fair, it works remarkably well for many common patterns. But “automatic” does not mean “perfect.” The default autoscaling mechanisms on most platforms are CPU-threshold-based, and they can struggle with highly dynamic or bursty workloads. When traffic spikes suddenly, the autoscaler may react too slowly, causing latency to spike. When traffic drops, it may scale down too aggressively, leaving too few warm instances for the next burst.

Research into smarter autoscaling has explored multi-metric approaches that consider not just CPU usage but also memory, request latency, queue depth, and congestion dynamics inspired by TCP slow-start algorithms. The idea is to ramp up capacity gradually when load increases, then back off as signals of congestion appear, similar to how network protocols regulate data flow.7International Journal of Cloud Applications and Computing. Adaptive Multi-Metric Autoscaling for Serverless Platforms: A TCP Slow-Start-Inspired Approach These techniques aim to reduce the SLA violations and scaling instability that simpler approaches produce under real-world conditions. For most applications, the default autoscaler is fine. For latency-sensitive services with unpredictable traffic, tuning the scaling behavior is part of the job.

Vendor Lock-in and the Open-Source Response

When you build on AWS Lambda, your function code is portable, but everything around it usually is not. The event triggers, the IAM permissions model, the logging integrations, the API gateway configuration, and the deployment tooling are all specific to AWS. Moving that application to Azure or Google Cloud means rewriting much of the glue that holds it together. This is vendor lock-in, and it is one of the most commonly cited concerns in the serverless world.

To address this, a number of open-source serverless platforms have emerged, including Knative, OpenFaaS, Apache OpenWhisk, and Fission. These let you run functions on your own Kubernetes clusters or on any cloud, preserving portability.8arXiv. Analyzing Open-Source Serverless Platforms: Characteristics and Performance The trade-off is that you are now managing infrastructure again, which defeats part of the point. You gain freedom from a single vendor but take on the operational burden that the managed platforms were supposed to eliminate. Whether that trade-off makes sense depends on your organization’s size, regulatory requirements, and tolerance for cloud-provider dependency.

A middle path that many enterprises are exploring is the hybrid model: running some workloads as serverless functions on a managed platform and others as containerized microservices on Kubernetes. This lets teams use serverless for event-driven, bursty tasks where the cost and scaling advantages are clearest, while keeping long-running or stateful services in containers where they have more control.9International Journal of Multidisciplinary Research in Science, Engineering and Technology. Microservices Architecture: Beyond Containerization vs. Serverless – A Hybrid Model for Enterprise Scale This hybrid approach is becoming the default architecture for large-scale applications rather than a pure all-in on either model.

Energy Efficiency and Environmental Impact

One of the less-discussed advantages of serverless computing is its energy footprint. Traditional cloud deployments often involve servers running around the clock, consuming power even when barely utilized. Serverless’s on-demand model means compute resources are allocated only when work needs to be done, and released immediately afterward.

Research suggests the savings can be substantial. One study found that serverless architectures reduced energy consumption by up to 70% and operational costs by up to 60% compared to traditional always-on deployments, because idle resource consumption was largely eliminated.10Sustainability. Reducing Environmental Impact with Sustainable Serverless Computing A separate evaluation focused on small-scale web applications found that serverless reduced idle-time energy consumption by up to 65% and overall energy usage by about 28% in low-traffic scenarios.11INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT. Evaluating the Energy Efficiency of Serverless Computing for Small-Scale Web Apps

These numbers make intuitive sense. If your web application gets a few hundred requests per hour, keeping a server running continuously to handle those requests wastes most of the electricity it draws. A serverless function that wakes up, processes each request, and goes back to sleep eliminates that waste. The environmental argument for serverless is strongest for exactly these low-to-moderate traffic patterns. High-traffic services that keep servers busy most of the time see a smaller relative benefit, because the baseline waste was already low.

Serverless for Scientific Computing and AI

Serverless started as a tool for web backends, but researchers have been pushing it into unexpected territory. In high-performance computing, the traditional model involves reserving large clusters of machines through a job scheduler, and those machines often sit partially idle between tasks. A software disaggregation approach using serverless functions allows idle cores and accelerators on allocated nodes to be utilized by placing functions on those underused resources, achieving near-native performance while squeezing more value from existing hardware.12PubMed Central. Software Resource Disaggregation for HPC with Serverless Computing

The funcX platform takes this further, offering what its creators call “serverless supercomputing.” It lets scientists register Python functions and execute them remotely on clusters, supercomputers, or cloud resources without worrying about the underlying scheduler or hardware. In demonstrations, funcX processed millions of functions across more than 65,000 concurrent workers.13arXiv. Serverless Supercomputing: High Performance Function as a Service for Science For researchers who need to run many independent computations, such as parameter sweeps, genomic analyses, or materials simulations, the serverless model lets them focus on the science rather than the cluster management.

AI inference is another frontier. Running a machine learning model to serve predictions looks a lot like a serverless workload: a request arrives, the model processes it, and a result goes back. But most models, especially large ones, need GPU access, and loading a model onto a GPU from scratch for each request would be painfully slow. One platform called Torpor addresses this by keeping models in main memory and swapping them onto GPUs only when requests arrive, using techniques like asynchronous redirection and pipelined execution to minimize the latency of that swap.14ACM Transactions on Architecture and Code Optimization. Enabling Low-Latency, GPU-Efficient Serverless Inference with Model Swapping Existing serverless platforms were not designed for GPU workloads, so this kind of work is essentially rebuilding the serverless model to accommodate hardware that doesn’t fit its original assumptions.

When Serverless Is the Wrong Choice

The enthusiasm around serverless sometimes obscures the cases where it genuinely does not make sense. Long-running processes hit platform-imposed time limits: most providers cap function execution at somewhere between five and fifteen minutes. If your workload involves training a machine learning model for hours or running a persistent WebSocket connection, you need a traditional server or container.

Workloads with predictable, steady traffic also lose the cost advantage. If your service runs at near-full capacity around the clock, the per-invocation billing of serverless can actually be more expensive than a reserved virtual machine doing the same work, because you are paying a premium for the elasticity you are not using.

Applications that require fine-grained control over the runtime environment, custom operating system configurations, or specific hardware may also find serverless too constraining. The abstraction that makes serverless easy is the same abstraction that takes away control. You cannot tune kernel parameters, install arbitrary system libraries, or choose your CPU architecture on most serverless platforms.

And debugging distributed serverless applications remains genuinely harder than debugging a monolithic one. When a request fans out across a dozen functions, each running independently and potentially on different machines, tracing the path of a single request through the system requires specialized observability tooling. The tools are improving, but the complexity is structural. Event-driven architectures, where functions respond to events that trigger other events in a cascade, can produce behaviors that are difficult to reason about and reproduce during debugging.

The Event-Driven Architecture Connection

Serverless computing and event-driven architecture are distinct concepts that work well together, and most real-world serverless applications are event-driven by nature. Rather than one service calling another directly and waiting for a response, components communicate by emitting and consuming events. A file lands in storage, which triggers a processing function, which writes a result to a database, which triggers a notification function. Each piece runs independently and asynchronously.

This combination improves responsiveness and fault tolerance. If the notification function is temporarily unavailable, the event can sit in a queue until it recovers, rather than causing the entire chain to fail. The components are loosely coupled, meaning you can update or replace one without touching the others.15The USA Journals TAJIIR. Serverless & Event-Driven Architectures: Redefining Distributed System Design For data-heavy applications that need to process streams of incoming information in near-real time, such as IoT telemetry, financial transactions, or log analysis, the serverless-plus-events combination is a natural fit because each event can be handled independently and at whatever scale the traffic demands.

The downside is cognitive. An event-driven system does not have a single call stack you can step through in a debugger. The flow of logic is distributed across event triggers, queues, and function invocations. Teams that adopt this pattern often find that getting the architecture right requires careful upfront design of event schemas and failure-handling strategies, because once events start flowing, the system’s behavior emerges from the interactions between components rather than from any one piece of code you can point to.