Platform engineer with 14 years building large-scale cloud and multi-tenant infrastructure and AI/ML systems. Deep experience leading secure, distributed services (Kubernetes, model serving, orchestration, and policy-as-code) across regulated enterprise environments while improving latency, reliability, and cost. Known for end-to-end ownership from infrastructure to AI workflows—sharded datastores, high-throughput search systems, and secure multi-cloud AI platforms. Focused on measurable impact such as availability improvements, faster deployment pipelines, and significant reductions in cloud spend and inference costs.

Jonathan Lin

Platform engineer with 14 years building large-scale cloud and multi-tenant infrastructure and AI/ML systems. Deep experience leading secure, distributed services (Kubernetes, model serving, orchestration, and policy-as-code) across regulated enterprise environments while improving latency, reliability, and cost. Known for end-to-end ownership from infrastructure to AI workflows—sharded datastores, high-throughput search systems, and secure multi-cloud AI platforms. Focused on measurable impact such as availability improvements, faster deployment pipelines, and significant reductions in cloud spend and inference costs.

Available to hire

Platform engineer with 14 years building large-scale cloud and multi-tenant infrastructure and AI/ML systems. Deep experience leading secure, distributed services (Kubernetes, model serving, orchestration, and policy-as-code) across regulated enterprise environments while improving latency, reliability, and cost.

Known for end-to-end ownership from infrastructure to AI workflows—sharded datastores, high-throughput search systems, and secure multi-cloud AI platforms. Focused on measurable impact such as availability improvements, faster deployment pipelines, and significant reductions in cloud spend and inference costs.

See more

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more

Language

Work Experience

Staff Software Engineer | AI/ML Infrastructure at In fields/platform software engineering role (employer not explicitly stated in resume text)
June 1, 2023 - June 1, 2026
Built a secure, multi-cloud AI/ML orchestration platform for regulated enterprise customers (SOC2, HIPAA). Architected a zero-trust control plane with per-tenant data planes across EKS/AKS/GKE/on-prem Kubernetes, achieving 99.9% uptime with tenant isolation via VPC peering and mTLS. Integrated Metaflow and Ray into a cloud-agnostic execution layer using DataBricks, Airflow, Snowflake, and SageMaker connectors, cutting model deployment lead time by 70%. Implemented append-only dataset and model lineage via MLflow registry with staged promotion/rollback, exposing SDK, REST, GraphQL APIs, and CLI for self-service. Scaled 500 V100s to thousands of inference QPS using TensorRT/ONNX runtime quantization, batching, and spot-based autoscaling, reducing cloud spend by 35% and inference cost by 30%. Embedded OPA policy-as-code, CVE scanning, Vault secrets, and prompt-injection guardrails into GitHub Actions, passing SOC2 with zero findings and catching malicious dependencies before production.
Staff Software Engineer | AI/ML Infrastructure at Indee d (platform team)
June 1, 2023 - June 1, 2026
Built a secure, multi-cloud AI/ML orchestration platform for regulated enterprise customers (SOC2, HIPAA). Architected a zero-trust control plane with per-tenant data planes across EKS/AKS/GKE and on-prem Kubernetes, achieving 99.9% uptime with tenant isolation via VPC peering and mTLS. Integrated MetaFlow and Ray into a cloud-agnostic execution layer using data bricks, Airflow, Snowflake, and SageMaker connectors, cutting model deployment lead time by ~70%. Implemented append-only dataset and model lineage with MLflow staged promotions and rollback; exposed Python SDK plus REST/GraphQL APIs and CLI. Scaled hundreds to thousands of V100s for inference QPS using TensorRT/ONNX runtime quantization, batching, and autoscaling, reducing cloud spend ~35% and inference cost ~30%. Embedded OPA Gatekeeper policies, CVE scanning, Vault secrets, and prompt-injection guardrails into GitHub Actions to pass SOC2 with zero findings. Onboarded dozens of data science teams to MetaFlow via examples, flo
Platform Software Engineer at Inferred (current/most recent employer not explicitly stated)
June 1, 2023 - June 1, 2026
Built sharded LSM storage supporting sub-31ms reads at extremely high search scale. Helped drive a Forge-to-GA migration achieving ~99.95% availability. Architected a secure multi-cloud AI platform that reduced model deployment time ~70%, cloud spend ~35%, and inference cost ~30%. Implemented deep Kubernetes-based (multi-tenant control/data-plane) design, model serving, and security-as-code using OPA policy-as-code and Vault, with mTLS and tenant isolation. Developed MLOps tooling for dataset/model lineage, staged promotion/rollback, and APIs/CLI for self-service access. Optimized inference using vLLM/TensorRT/ONNX-style runtime strategies including quantization, batching, and autoscaling. Performed reliability work including incident response readiness and disaster recovery practices; added chaos testing and automated security scanning gates to CI/CD.
Senior Software Engineer at Atlassian (Forge)
October 1, 2019 - June 1, 2023
Helped take Forge—Atlassian’s serverless app platform for Jira and Confluence—from early development to May 2021 general availability. Built React UI kit issue panels, Confluence macro components, and Node.js resolver functions on AWS Lambda, enabling ~25,000 third-party developers to ship hosted apps without managing servers. Designed idempotent Forge event handlers, web triggers, scheduled jobs, and async queues so long-running work continues independently of request paths. Increased throughput by ~50% using Lambda provisioned concurrency, KVS caching, Jira REST batching, and Runtime 2.0 while maintaining 99.95% availability and 200–500ms latency. Built phased Connect-to-Forge migration tooling across registration, OAuth 2.0 authorization, and Forge Storage data moves, enabling partner apps to shift with zero downtime as the beta reached ~3,500 apps.
Senior Software Engineer at Atlassian
October 1, 2019 - June 1, 2023
Helped take Forge (Atlassian’s serverless app platform for Jira/Confluence) from early development through May 2021 general availability. Built UI kit issue panels, Confluence macros, and resolver functions on AWS Lambda enabling third-party developers to host apps without servers. Implemented idempotent event handlers, web triggers, scheduled jobs, and async queues for long-running work. Improved throughput substantially using provisioned concurrency, caching, batch handling, and runtime tuning to maintain very high availability and low latency. Led staged migration tooling for Connect-to-Forge across registration, OAuth, and data migration with zero downtime as partner app adoption scaled.
Senior Backend Engineer at Inferred job within job-search / recommendations platform (employer not explicitly stated in resume text)
June 1, 2012 - July 1, 2019
Built core search, storage, ingestion, and recommendation-serving systems for a job-search backend handling a billion-plus queries per month across 50+ countries. Designed a sharded, replicated LSM store with Lucene full-text indexes, migrating off MySQL via zero-downtime dual writes and supporting ~31 ms reads at billion-scale monthly searches. Rebuilt job ingestion as idempotent RabbitMQ workers with exponential backoff and rate limiting, absorbing ~35M postings per day without slowing customer-facing search. Built a hybrid Hadoop + Memcached recommendation layer using incremental LSM model segments, growing recommendation clicks from 3% to 14% with ~30% A/B lift. Ran production readiness and PagerDuty on-call across three data centers using JMeter load tests and staged rollouts to keep sub-50 ms lookups during spikes.
Senior Backend Engineer at In d eed (job search backend systems)
June 1, 2012 - July 1, 2019
Built core search, storage, ingestion, and recommendation services for a job-search backend handling over a billion-plus queries per month across 50+ countries. Designed a sharded, replicated LSM store with Lucene full-text indexes; migrated off MySQL through zero-downtime dual writes and served ~31ms reads at billion-plus monthly searches. Rebuilt job ingestion as idempotent RabbitMQ workers with exponential backoff and rate limiting, absorbing ~35M postings per day without slowing customer-facing search. Implemented a hybrid Hadoop + cached recommendation serving layer using incremental LSM model segments, growing recommendation click-through from 3% to 14% with ~30% A/B lift. Ran production readiness and PagerDuty on-call for search across three data centers using JMeter load tests and staged rollouts to keep sub-50ms lookups during spikes.
Senior Backend Engineer at Inferred (company not explicitly stated)
June 1, 2012 - July 1, 2019
Built core job search backend components including search, storage, ingestion, and recommendation serving for a job-search system handling over a billion queries per month across 50+ countries. Designed and implemented a sharded, replicated LSM store with Lucene full-text indexes and migrated off MySQL using dual writes and zero-downtime cutovers to achieve ~31ms reads at very high scale. Rebuilt job ingestion as idempotent RabbitMQ workers with backoff/rate limiting to absorb extremely high daily posting volume without slowing customer-facing search. Implemented a hybrid Hadoop + cached recommendation-serving layer based on incremental LSM model segments, increasing recommendation clicks materially through A/B testing. Ran production readiness and on-call responsibilities, using load tests and staged rollouts to hold sub-50ms lookups during traffic spikes.

Education

Bachelor of Science in Computer Science at The University of Texas at Austin
January 1, 2007 - January 1, 2012
Bachelor of Science in Computer Science at The University of Texas at Austin
January 1, 2007 - January 1, 2012
Bachelor of Science in Computer Science at The University of Texas at Austin
January 1, 2007 - January 1, 2012

Qualifications

Industry Experience

Software & Internet, Healthcare, Financial Services, Professional Services

Experience Level

Expert
Expert
Expert
Expert
Expert
Expert
Expert
Intermediate
See more