Job Overview
We’re looking for a Cloud & Infrastructure Engineer to build, secure, and maintain scalable platform infrastructure and delivery pipelines. The ideal candidate combines hands-on experience with managed PaaS platforms and core AWS services (RDS, S3, Lambda) with strong capabilities in Infrastructure as Code, automated CI/CD, and transactional email deliverability. This role will collaborate closely with project leads and software engineers, establish comprehensive observability and security controls, and deliver reliable, automated, and well-documented cloud environments.
Core Tasks
1. Cloud architecture and platform
• Assess the current infrastructure — what is running, how it is configured, what is manual, and where the risk sits — and produce a written assessment with prioritized findings.
• Define and implement the target infrastructure across PaaS platform, e.g., Heroku, Render, Elastic Beanstalk, App Runner and supporting AWS services, with environment separation for dev / staging / production.
• Define infrastructure as code so environments are reproducible rather than hand-configured. Terraform / CloudFormation / CDK / Pulumi, as agreed.
• Implement IAM roles and least-privilege access, secrets management, and network configuration (VPC, security groups, TLS).
2. Data and storage layer
• RDS: provision and tune the database — instance sizing, parameter groups, connection pooling, read replicas if warranted — and configure automated backups, point-in-time recovery, and encryption at rest.
• Run and document a restore test. A backup that has never been restored is not a backup.
• S3: configure buckets for uploads / assets / exports / logs with correct access policies, lifecycle rules, versioning, and encryption. No public buckets that should not be public.
• Set up CDN delivery for static and user-uploaded assets where appropriate.
3.Serverless and background processing
• Lambda: build and deploy the functions in scope — e.g., scheduled jobs, event-driven processing, webhook handlers, image or file processing — with appropriate triggers, concurrency limits, and timeouts.
• Handle failure properly: retries, dead-letter queues, idempotency, and alerting on repeated failure.
• Bring Lambda deployments into the same CI/CD pipeline and IaC definitions as the rest of the stack, not a separate manual path.
4.Email and outbound messaging
• SMTP services: configure transactional email through SES / SendGrid / Postmark / other for e.g., account, notification, and system emails.
• Set up domain authentication — SPF, DKIM, DMARC — and manage sending reputation, warmup, and dedicated IP if warranted.
• Configure bounce, complaint, and suppression handling, plus deliverability monitoring so a drop in inbox placement is visible before users report it.
5. CI/CD and release engineering
• Build pipelines in GitHub Actions / GitLab CI / CircleCI / other that run tests, lint, build, and deploy on merge, with environment promotion from staging to production.
• Implement deployment strategy and rollback: blue-green / rolling / canary, as agreed, with a rollback that is a single documented action rather than an improvisation.
• Automate database migrations as part of deploy, safely and reversibly.
• Reduce pipeline runtime so the team is not waiting on it.
6. Observability, security, and cost
• Set up centralized logging, metrics, and alerting through CloudWatch / Datadog / other, with alerts that page on real problems and stay quiet otherwise.
• Define and monitor the health indicators that matter for the team: uptime, error rate, latency, queue depth, job failure.
• Review the stack against HIPAA / PCI / SOC 2 / other requirements where applicable, and document what is in place and what is not.
• Establish cost visibility and tagging; identify and remove waste.
Deliverables
• Infrastructure assessment with prioritized findings, reviewed with project lead.
• Architecture diagram of the target infrastructure, with data flow and trust boundaries.
• Infrastructure as code covering all environments, in repository.
• Configured and hardened RDS, S3, Lambda, and PaaS environments, with a documented and tested restore procedure.
• Transactional email configured and authenticated, with deliverability monitoring in place.
• CI/CD pipelines deploying dev / staging / production with automated tests, migrations, and one-step rollback.
• Monitoring, logging, and alerting, with an on-call runbook covering the top failure modes and what to do about each.
• Cost baseline and tagging scheme, with identified savings.
• Handoff package: infrastructure documentation, access inventory, known limitations, recommended next steps, and a walkthrough with the team.
Must Haves
• 3+ years of engineering experience
• PaaS — production experience running applications on a managed platform, and clear judgment about where the platform's abstractions stop being enough.
• AWS RDS — provisioning, tuning, backup and recovery, and diagnosing the database problems that surface as application slowness.
• AWS S3 — bucket policies, lifecycle management, encryption, and secure handling of user-uploaded content.
• AWS Lambda — building, deploying, and operating serverless functions, including their less obvious failure modes.
• SMTP services — configuring transactional email, domain authentication (SPF, DKIM, DMARC), and deliverability management.
• CI/CD — you have built pipelines a team actually relied on, and have made deploying boring for people who were previously nervous about it.
Nice to Haves
• Infrastructure as code at depth in Terraform / CloudFormation / CDK.
• Containers and orchestration where the stack calls for it.
• Experience with HIPAA / PCI / SOC 2 compliance in a cloud environment.
• Domain experience in medical software and/or e-commerce.
• Prior contract work where you owned infrastructure end to end and handed it off cleanly.
Originally posted on Himalayas