Linux containers in 500 lines of code

(blog.lizzie.io)

64 points | by mkornaukhov 6 hours ago

5 comments

  • abidinberkay 26 minutes ago
    This was written in 2016. What would be different if you wrote it today? For example would cgroups v2 or newer seccomp features change that much?
  • js2 2 hours ago
    (2016). Previous submissions w/comments:

    https://news.ycombinator.com/item?id=30623372 (250 points | March 10, 2022 | 27 comments)

    https://news.ycombinator.com/item?id=22232705 (267 points | Feb 4, 2020 | 29 comments)

    https://news.ycombinator.com/item?id=15608435 (440 points | Nov 2, 2017 | 53 comments)

  • setheron 2 hours ago
    I have written https://fzakaria.com/2020/05/31/containers-from-first-princi... a while ago in similar vein.
  • ranger_danger 2 hours ago
    > I wanted specifically to find a minimal set of restrictions to run untrusted code.

    I don't think we should consider containers to be a security boundary. Even full VMs can be escaped, and have been, many times.

    The fact that this is possible in the first place makes me think we need a much better approach.

    • bityard 42 minutes ago
      Depends on the threat model. Security is not black-and-white.

      Containers protect against "I don't trust this curlpipe to not crap all over my dotfiles," rather than, "there might be a sandbox escape attack in this random file I downloaded."

      If a VM is not sufficient for your threat model, I'm curious what is?

    • chubot 2 hours ago
      As far as I know, Firecracker, gVisor, and Kata Containers are the solution here. They use VM primitives (x64_64 and ARM64 extensions) and have lighter codebases

      https://firecracker-microvm.github.io/

      https://gvisor.dev/

      https://katacontainers.io/

      But I don't have any direct experience with any of them. I'd be curious what people who have built on top of them think

      edit: OK it looks like Kata can use Firecracker, so as far as isolation, it's either Firecracker or gVisor. And Firecracker is the VMM I mentioned, but gVisor is quite different -- it's more like a user space kernel that emulates syscalls.

      • johnsmith1840 10 minutes ago
        I've deploy gvisor, done basic test of firecracker and an honest attempt at production kata.

        Firecracker and gvisor are nice systems not horrible to use, gvisor isn't quite the same security level though.

        Kata is HARD to make. The technical know how to make that in production is awe inspiring. I wanted to use it but it was so complicated to integrate into a cluster I literally just gave up and mirrored raw VMs into the cluster which was alot easier actually.

        Kata also breaks any potential of confidential VM unless you're a virtualization wizard.

        You should go check out redhat's confidential container method for a production design overview. Their ARO self hosted system.

      • binsquare 1 hour ago
        I'm going to toss in smolvm as well because firecracker needs some expertise to make the box usable and secure.

        https://github.com/smol-machines/smolvm

      • laurencerowe 1 hour ago
        As I understand it Kata supports multiple VMM backends, Firecracker, QEmu, Cloud Hypervisor, and their own Dragonball. Except QEmu, I believe those are all built on crates in the rust-vmm ecosystem, each making slightly different tradeoffs.
    • stryan 50 minutes ago
      Podman supports using KVM backed virtualization for containers via libkrun: `podman run --runtime=krun` . Still not the end-all-be-all security boundary, but better I think.
    • raesene9 1 hour ago
      I definitely wouldn't trust standard Linux style containers that expose a shared Linux kernel at the moment, there's been far too many LPE and container breakout vulnerabilities this year. It's possible that in future if the kernel gets a lot more hardened, that could change but things like Firecracker are a better bet from a security standpoint.
    • mdspan 37 minutes ago
      I think until something hardware-based like CHERI becomes widely deployed (which seems extremely unlikely in the near to mid term given), we're going to keep seeing VM escape CVEs pop up indefinitely.
  • tankiya 2 hours ago
    [flagged]