惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threatpost
S
Schneier on Security
P
Palo Alto Networks Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
S
Securelist
T
Threat Research - Cisco Blogs
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
V
Vulnerabilities – Threatpost
AI
AI
C
CERT Recently Published Vulnerability Notes
C
Cyber Attacks, Cyber Crime and Cyber Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Know Your Adversary
Know Your Adversary
AWS News Blog
AWS News Blog
TaoSecurity Blog
TaoSecurity Blog
O
OpenAI News
Cyberwarzone
Cyberwarzone
G
GRAHAM CLULEY
SecWiki News
SecWiki News
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Simon Willison's Weblog
Simon Willison's Weblog
I
Intezer
C
CXSECURITY Database RSS Feed - CXSecurity.com
C
Cisco Blogs
K
Kaspersky official blog
Spread Privacy
Spread Privacy
S
Security @ Cisco Blogs
Hacker News - Newest:
Hacker News - Newest: "LLM"
IT之家
IT之家
有赞技术团队
有赞技术团队
B
Blog
T
Tailwind CSS Blog
PCI Perspectives
PCI Perspectives
P
Privacy & Cybersecurity Law Blog
Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Webroot Blog
Webroot Blog
博客园 - 叶小钗
Cisco Talos Blog
Cisco Talos Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
H
Hacker News: Front Page
The Cloudflare Blog
C
Cybersecurity and Infrastructure Security Agency CISA
美团技术团队
V
V2EX
腾讯CDC
S
SegmentFault 最新的问题
Security Archives - TechRepublic
Security Archives - TechRepublic

Ivan on Containers, Kubernetes, and Server-Side

A grounded take on agentic coding for production environments Server-Side Playgrounds Reimagined: Build, Boot, and Network Your Own Virtual Labs [not a] Kubernetes 101 - Pods, Deployments, and Services As an Attempt To Automate Age-Old Infra Patterns JavaScript or TypeScript? How To Benefit From the Dichotomy On Software Design... and Good Writing Building a Firecracker-Powered Course Platform To Learn Docker and Kubernetes How To Publish a Port of a Running Container What Actually Happens When You Publish a Container Port A Visual Guide to SSH Tunnels: Local and Remote Port Forwarding Debugging Containers Like a Pro Docker: How To Debug Distroless And Slim Containers How To Extract Container Image Filesystem Using Docker | iximiuz Labs In Pursuit of Better Container Images: Alpine, Distroless, Apko, Chisel, DockerSlim, oh my! How To Start Programming In Go: Advice For Fellow DevOps Engineers Kubernetes Ephemeral Containers and kubectl debug Command How To Develop Kubernetes CLIs Like a Pro Docker Container Commands Explained: Understand, Don't Memorize | iximiuz Labs Learning Docker with Docker - Toying With DinD For Fun And Profit How To Extend Kubernetes API - Kubernetes vs. Django The Influence of Plumbing on Programming How To Call Kubernetes API from Go - Types and Common Machinery How To Call Kubernetes API using Simple HTTP Client Kubernetes API Basics - Resources, Kinds, and Objects OpenFaaS - Run Containerized Functions On Your Own Terms Learning Containers From The Bottom Up Docker Containers vs. Kubernetes Pods - Taking a Deeper Look | iximiuz Labs Learn-by-Doing Platforms for Dev, DevOps, and SRE Folks How HTTP Keep-Alive can cause TCP race condition How to Work with Container Images Using ctr | iximiuz Labs Multiple Containers, Same Port, no Reverse Proxy... Exploring Go net/http Package - On How Not To Set Socket Options Disposable Local Development Environments with Vagrant, Docker, and Arkade My Choice of Programming Languages Prometheus Is Not a TSDB How to learn PromQL with Prometheus Playground Prometheus Cheat Sheet - Basics (Metrics, Labels, Time Series, Scraping) Rust - Writing Parsers With nom Parser Combinator Framework pq - parse and query log files as time series Prometheus Cheat Sheet - Moving Average, Max, Min, etc (Aggregation Over Time) Prometheus Cheat Sheet - How to Join Multiple Metrics (Vector Matching) The Need For Slimmer Containers Understanding Rust Privacy and Visibility Model Bridge vs. Switch: Takeaways from a Real Data Center Tour | iximiuz Labs From LAN to VXLAN: Networking Basics for Non-Network Engineers | iximiuz Labs KiND - How I Wasted a Day Loading Local Docker Images Go, HTTP handlers, panic, and deadlocks Exploring Kubernetes Operator Pattern Making Sense Out Of Cloud Native Buzz Service Discovery in Kubernetes: Combining the Best of Two Worlds API Developers Never REST How Container Networking Works: Building a Bridge Network From Scratch | iximiuz Labs Traefik: canary deployments with weighted load balancing Service Proxy, Pod, Sidecar, oh my! You Need Containers To Build Images You Don't Need an Image To Run a Container Not Every Container Has an Operating System Inside Working with container images in Go Master Go While Learning Containers Implementing Container Runtime Shim: Interactive Containers How to use Flask with gevent (uWSGI and Gunicorn editions) My 10 Years of Programming Experience Implementing Container Runtime Shim: First Code Implementing Container Runtime Shim: runc Kubernetes Repository On Flame Dealing with process termination in Linux (with Rust examples) conman - [the] Container Manager: Inception Journey From Containerization To Orchestration And Beyond Linux PTY - How docker attach and docker exec Commands Work Inside Illustrated introduction to Linux iptables From Docker Container to Bootable Linux Disk Image Пишем свой веб-сервер на Python: протокол HTTP 9001 способ создать веб-сервер на Python Explaining async/await in 200 lines of code Explaining event loop in 100 lines of code Save the day with gevent Пишем свой веб-сервер на Python: процессы, потоки и асинхронный I/O Truly optional scalar types in protobuf3 (with Go examples) Node.js Writable streams distilled Node.js Readable streams distilled How to on starting processes (mostly in Linux) Дайджест интересных ссылок – Июль 2016 Пишем свой веб-сервер на Python: сокеты Наследование в JavaScript Мастерить!
DevOps, SRE, and Platform Engineering
Ivan Velichko · 2021-08-01 · via Ivan on Containers, Kubernetes, and Server-Side

I compiled this thread on Twitter, and all of a sudden, it got quite some attention. So here, I'll try to elaborate on the topic a bit more. Maybe it would be helpful for someone trying to make a career decision or just improve general understanding of the most hyped titles in the industry.

DevOps, SRE, and Platform Engineering (thread)

Sharing my understanding of things after working in this domain for about two years.

Starting from the clearest one.

Dev - this is about application development, aka business logic. The only one that makes money for a company.

— Ivan Velichko (@iximiuz) July 31, 2021

Level up your server-side game — join 20,000 engineers getting insightful learning materials straight to their inbox.

During my career, I used to work in teams and companies where as a developer, I would push code to a repository and just hope that it would work well when some mythical system administrator would eventually take it to production. I also was in setups where I would need to provision bare-metal servers on Monday, figure out the deployment strategy on Tuesday, write some business logic on Wednesday, roll it out myself on Thursday, and firefight a production incident on Friday. And all this without even being aware of the existence of fancy titles like DevOps or SRE engineer.

But then people around me started talking DevOps and SRE, comparing them with each other, and compiling awesome lists of resources. New job opportunities began emerging, and I quickly jumped into the SRE train. So, below is my experience of being involved in all things SRE and Platform Engineering from the former Software Developer standpoint. And yeah, I think it's applicable primarily for companies where the product is some sort of a web-facing service. This is the kind of company I spent ten years working for. People doing embedded software or implementing databases probably live in totally different realities.

What is Development

This one is the simplest to explain. Development - is about application programming, i.e., writing the business logic of your main product. This is the only activity among the three ones being discussed here that directly makes money for the company.

The only one that makes for a company is of course sales, everything else is expenditure :)

— ⊥nɹqöSןɐʌıʞ (@slavadotcom) July 31, 2021

IMO, development is super hot! As a developer, you quickly start thinking that you are the most important person around. Without your code, there is nothing. But apparently, just writing code often isn't enough. The code needs to be delivered to production and executed there.

I'd been carrying the Software Developer (or Software Engineer) title since the very beginning of my career in 2011. And I still remember the pain quite vividly - I always wished to have control over deploying my code. And I rarely had it. Instead, there would be some obscure procedure when someone, usually not even your senior colleague, would have access to production servers and deploy the code there for you. So, if after pushing the changes to the repository, you got unlucky enough to notice a bug only on the live version of your service, you'd need to beg for an extra rollout. It most definitely sucked.

What is DevOps

I'll not even try to quote the official definition here. Instead, I'll share the first-hand experience. For me, DevOps was a cultural shift giving development teams more control over shipping code to production. The implementation could vary. I've been in setups where developers would just have sudo on production servers. But probably the most common approach is to provide development teams with some sort of CI/CD pipelines.

In an ideal GitOps world, developers would still be just pushing code to repositories. However, there would be a magical button somewhere at the team's disposal that would put the new version on live or maybe even provision a new piece of infrastructure to cover the new requirements.

The original idea of DevOps is probably much broader than just that. But from what I see in the job descriptions, what I hear from recruiters trying to hunt me for a DevOps position, and what I managed to gather from my fellow colleagues carrying the DevOps engineer title, most of the time, it's about creating an efficient way to deploy stuff produced by Development. In more advanced setups, DevOps may also be concerned with other things improving the Development velocity. But DevOps itself is never concerned with the actual application business logic.

What is SRE

There is a excellent series of books by Google explaining the idea of the Site Reliability Engineering and, what's even more important for me, sharing some real tech practices conducted by Google SREs. In particular, it says that SRE is just one of the ways to implement the DevOps culture - class SRE implements DevOps {}.

This explanation didn't really help me much. But what was even more puzzling, subconsciously, I always felt excited while reading SRE job descriptions and got bored quickly by the DevOps ones... So, there was clearly a difference but, for a long time, I couldn't distill it.

Of course, that's just about my personal preferences, but whenever someone mentions configuring a CI/CD pipeline, I always got depressed. And the DevOps job descriptions nowadays are full of such responsibilities. Don't get me wrong, CI/CD pipelines are amazing! I'm always glad when I have a chance to use one. But setting them up isn't a thing I enjoy the most. On the contrary, when someone asks me to jump in and take a look at a bleeding production, be it chasing a bug, a memory leak, or performance degradation, I'm always more than just happy to help.

Developing code and shipping it to production still doesn't give you the full picture. Someone needs to keep the production alive and healthy! And that's how I see the place of SRE in my model of the world.

Google's SRE book focuses on monitoring and alerting, defining SLOs of your services and tracking error budgets, incident response and postmortems. These are the things one would need to apply to make the production reliable. Facebook has a famous Production Engineer role, but it's pretty hard to distinguish it from a typical SRE role, judging only by the job description.

Here is also a great tweet that kind of confirms my feeling that the primary focus of SRE is production.

My very simplified answer when someone says what is the difference between SRE and DevOps.

* SRE = focused primarily on production
* DevOps = focused primarily on CI/CD and developer velocity

— Tammy Bryant Butow ⚓ (@tambryantbutow) June 16, 2021

And one more:

I like it! My typical answer is:
SRE works from Production backward. DevOps works from development forward. Somewhere in the middle, they meet.

— Gary Pochron (@garypochron) June 16, 2021

So, DevOps keeps production fresh. SRE keeps production healthy.

UPD: Bruce Dominguez recently published an article with some thoughts on what makes SRE teams different from Ops or Platform Engineering teams, and it echoed my experience. The ROAD to SRE describes the key focus areas of an SRE team (and IMO, the order matters):

  • Response - setting up efficient incident response culture
  • Observability - instrumenting, monitoring, and alerting
  • Availability & Reliability - SLI/SLOs and failure management
  • Delivery - efficient building, provisioning, and deployment (IaC, CI/CD, etc.)

UPD 2: A recent write-up SRE vs. DevOps: What’s the Difference Between Them? by spacelift.io folks reinforces my take on the topic. To a large extent, the information in the article repeats my personal experience.

What is Platform Engineering

When I used to be the only engineer in a startup, a substantial part of my job was to turn some generic resources I'd rent from the infrastructure provider into something more tailored for the company's needs. So, I had a bunch of scripts to provision a new server, some understanding of how to provide network connectivity between our servers in different data centers, some skills to replicate the production setup on staging, and maybe even write one or two daemons to help me with log collection. I didn't really understand it, but these things constituted our Platform.

Joining a much bigger company and starting consuming infra-related resources brought me to a realization that there is a third area of focus that might be quite close to DevOps and SRE. It's called Platform Engineering.

From my understanding, Platform Engineering focuses on developing an ecosystem that can be efficiently used from the Dev, Ops, and SRE standpoints.

There might be quite some code writing in Platform Engineering. Or, it could be mostly about configuring outsourced/third-party things. But again, it's not about the primary business logic of your product - it's about making some basic infrastructure more suitable for the day-to-day needs.

Platform Engineering is not about infrastructure development.

Platform Engineering is about enabling others to do whatever they want to do.

Twitter was a magnificent platform in the early days, nothing to do with infra. https://t.co/6hKdaIxYOt

— Ivan Pedrazas (@ipedrazas) July 31, 2021

To be honest, I don't see a contradiction between my way of seeing Platform Engineering and the explanation from this tweet. Development needs infrastructure to run the code. So, if Platform Engineering is about enabling others to do whatever they want to do, at least in part, it should be concerned with infrastructure development.

I have a feeling that in a bigger setup, when a company would have thousands of bare-metal servers in its own data centers, a Platform Engineering might start from managing this fleet of machines. So, some sort of inventory software might need to be installed or even developed internally. Installing operating systems and basic packages on the servers being provisioned would probably also fall into the Platform Engineering responsibility. But the cross-cutting theme behind all these tasks would be providing some higher-level self-service building blocks for dev teams.

Luckily, clouds, IaaS, and SaaS made Platform Engineering operating on much higher layers. All the basic fleet management tasks are already solved for you. And even orchestration of your workloads is solved by projects like Kubernetes or (rather) AWS ECS. However, the solution is quite generic, while your teams are likely to deploy pretty similar microservices. So, providing them with a default project template that would be integrated with the company's metrics and logs collection subsystems would make things moving much faster.

UPD 3: The Future of Ops Is Platform Engineering by Charity Majors is a much more informed and versatile take on the topic. Do recommend for thorough reading!

What's about titles?

So far, I was deliberately avoiding talking about roles and titles. Development, Operations, SRE, and Platform Engineering for me are about areas of focus. And to a much lesser extent about titles. One person can be a Dev this week, then an Ops on the next week, and an SRE on the week after.

From my experience, the separation between Dev, Ops, SRE, and PE becomes more apparent when the company size gets bigger. A bigger company size usually means more specialists and fewer generalists. That's how you end up with dedicated SRE teams and a Platform Engineering department. But of course, it's not a strict rule. For instance, with my SRE title, I spent like a year doing all things true SRE (SLO, monitoring, alerting, incident response) and then transitioned into Platform Engineering, where I do more infra development than traditional SRE. YMMV.

Where Security goes?

When my thread went viral on Twitter, someone asked me how security fits into this picture. And that's indeed a very good question! But I don't have a simple answer. For me, a reasonable approach is to make security a cross-cutting theme in all Dev, Ops, SRE, and PE. Different security concerns can be addressed on different layers using different tools. For instance, Development could be concerned with preventing SQL injections while Platform folks could harden the networking by configuring some fancy cilium policies.

Instead of conclusion

Don't forget, all the things above are IMO 😉

Level up your server-side game — join 20,000 engineers getting insightful learning materials straight to their inbox: