Service Configuration
Customize how your services run in the cloud using infrastructure.json.
Overview
The .tsdevstack/infrastructure.json file lets you override default settings per environment:
Creating infrastructure.json
If you don't have an infrastructure.json file yet, create it in your .tsdevstack directory:
Then add the minimum required content:
IDE Autocomplete
The $schema property enables intelligent autocomplete and validation in VS Code and other editors that support JSON Schema.
What you get:
- Autocomplete for all configuration options (CPU, memory, database tiers, etc.)
- Inline validation for invalid values
- Hover documentation for each property
How it works:
- Add
"$schema": "./infrastructure.schema.json"as the first property in yourinfrastructure.json - The framework includes
infrastructure.schema.jsonin the.tsdevstackdirectory - Your IDE automatically provides suggestions and validation
If autocomplete isn't working, ensure the infrastructure.schema.json file exists in your .tsdevstack directory. Run npx tsdevstack infra:init --env <environment> to generate it.
Service Options
The JSON schema is provider-aware — IDE autocomplete will only show values valid for your cloud provider.
Scaling Configuration
minInstances
Controls the minimum number of instances always running.
Scale-to-zero is a platform feature of Cloud Run and Container Apps, and cold start times vary:
On AWS every ECS service (NestJS services, Next.js frontends, workers and Kong) runs at least one task. minInstances: 0 fails validation in infra:generate, infra:plan, infra:deploy and infra:status, with a hint to set 1 or more. Backend services on AWS are reachable only inside the VPC, so there is no public entry point that could wake a stopped service. See AWS Cost Estimation for what always-on costs.
Recommendations:
- Dev on GCP or Azure:
0- Save costs, cold starts are acceptable - Prod critical paths:
1+- Avoid cold starts for user-facing APIs - Detached workers:
1+- Workers must always be running to poll Redis queues. On GCP and Azure the framework raises0to 1 for workers; on AWS it fails validation like any other service
maxInstances
Limits how many instances can run during high traffic.
- Higher values handle more concurrent users
- Each instance costs money while running
- The container runtime auto-scales based on CPU utilization and request queue. On AWS, NestJS services, Next.js frontends, workers and Kong scale on CPU (target tracking at 70%) between
minInstancesandmaxInstances
Example: If each instance handles 80 concurrent requests and you expect 800 peak concurrent users, set maxInstances: 10 minimum.
Resource Configuration
CPU
Memory
GCP has the widest range (256Mi–8Gi). AWS supports larger instances (up to 16Gi). Azure is the most constrained (0.5Gi–4Gi, no sub-Gi values).
Kong Gateway
Configure the API gateway separately:
Kong defaults:
Keep Kong at minInstances: 1 or more: every API request goes through it. On AWS, 0 fails validation.
Detached Workers
Configure BullMQ worker scaling and resources:
Workers support the same options as services (minInstances, maxInstances, cpu, memory). The framework enforces minInstances >= 1 because workers must always be running to poll Redis queues: on GCP and Azure 0 is overridden to 1, on AWS it fails validation.
On AWS, workers get auto-scaling resources (CPU-based target tracking at 70%), like services and Kong. On GCP and Azure, the container runtime handles scaling natively.
Changing a worker's maxInstances affects the database connection pool calculation. After changing scaling config, redeploy all services and workers to rebalance pool sizes.
Database
Configure managed PostgreSQL:
Database tiers by provider:
GCP (Cloud SQL):
db-f1-micro- Development (shared CPU)db-g1-small- Small productiondb-n1-standard-1- Standard production (1 vCPU)db-n1-standard-2- Larger production (2 vCPU)db-n1-standard-4- High-traffic production (4 vCPU)
AWS (RDS PostgreSQL):
db.t3.micro- Development (~$15/mo)db.t3.small- Small productiondb.r6g.large- Standard productiondb.r6g.xlarge- High-traffic production
Azure (PostgreSQL Flexible Server):
B_Standard_B1ms- Development (burstable, ~$14/mo)B_Standard_B2s- Small production (burstable)GP_Standard_D2s_v3- Standard production (general purpose)GP_Standard_D4s_v3- High-traffic production (general purpose)
Custom server name (Azure)
On Azure the PostgreSQL Flexible Server name is a global DNS name (<name>.postgres.database.azure.com), so it has to be unique across all of Azure, not only in your subscription. tsdevstack names it {project}-{env}-postgres. If someone else already holds that name, the deploy fails with ServerNameAlreadyExists. Typical causes: another project uses the same project name, or a server with that name still exists in a subscription you deleted.
Pick a different name for that environment with serverName:
Rules: 3 to 63 characters, lowercase letters, digits and hyphens, no hyphen at the start or the end. Set it per environment; environments without it keep the default name. Your services' connection strings follow on their own, since the deploy builds them from the server address Terraform reports.
The field is Azure only. On GCP and AWS the schema flags it in your editor and infra:status reports it as an error.
Set serverName before the first deploy of an environment. Changing it on an environment that is already running makes Terraform replace the server: the old server and every database on it are destroyed and a new, empty one is created. If you really need to rename a live server, back up first (pg_dump) and check npx tsdevstack infra:plan before you deploy.
Redis
Configure managed Redis:
GCP (Memorystore):
BASIC- Single instance, no failoverSTANDARD_HA- High availability with automatic failover
AWS (ElastiCache):
cache.t3.micro- Development (~$12/mo)cache.t3.small- Small productioncache.r6g.large- Standard production
Azure (Managed Redis):
Balanced_B0- Development (~$13/mo, clustered)Balanced_B1- Small productionBalanced_B3- Standard production
Azure Managed Redis uses EnterpriseCluster policy. BullMQ requires a {bull} prefix to avoid CROSSSLOT errors — see Azure Architecture for details.
Frontend Domains
Configure domains for frontend services:
Load Balancer
Configure domain redirects and API domain:
When you own multiple domains (e.g., .com, .io, .app), add alternates to redirectDomains. All traffic to those domains redirects to your canonical domain set in the DOMAIN secret.
Access Control
Keep non-production environments out of search engines:
Environment password protection was removed. The accessControl.protected and accessControl.cookieTtlHours fields are no longer accepted, and a config that still sets them fails validation. Remove them from infrastructure.json.
Scheduled Jobs
Configure cron-based scheduled jobs in the scheduledJobs array. See the Scheduled Jobs guide for full configuration, service-side implementation, authentication, and provider architecture.
With the auth template, every environment needs the sync-api-key-usage job; infra:generate warns and prints the entry when it is missing. See API Keys.
WAF Rules
The framework includes default WAF rules for common attacks (SQL injection, XSS, path traversal, command injection, SSRF, scanner fingerprints, and more). You can customize the global rate limit and add custom rules per provider.
Global rate limit
Override the default rate limit (1000 requests per 60 seconds per IP) across all providers:
The framework translates this to each provider's native format:
- GCP: Direct pass-through to Cloud Armor throttle rule
- AWS: Scaled to 5-minute window (AWS minimum). 1000/60s becomes 5000 per 5 minutes
- Azure: Scaled to per-minute buckets. 500/30s becomes 1000 per 1 minute
Custom rules (GCP)
GCP uses CEL expressions for custom rules:
Custom rules (AWS)
AWS uses statement-based rules with byte match, rate-based, and geo match types:
Custom rules (Azure)
Azure uses match conditions with operators and transforms:
Match condition options:
Schema validation
The infrastructure.schema.json is provider-aware — it only shows the custom rule format relevant to your cloud provider. IDE autocomplete will guide you to the correct syntax.
Upload Size Limits
Configure the maximum traditional upload size through Kong:
This sets Kong's client_max_body_size. Requests exceeding this limit receive a 413 Content Too Large response.
Hard ceilings per provider (cannot be exceeded even with higher configuration):
For files larger than these limits, use presigned URL uploads which bypass the WAF/LB/Kong chain entirely. See Object Storage for details.
Cost Optimization
Development environments
Production environments
Applying Changes
After modifying infrastructure.json:
Changes are applied on the next deployment. Some changes (like database tier) may cause brief downtime.