Meta tags:
description=
A blog about distributed systems, container orchestration platforms, and AI infrastructure
;
author= Anton Kirillov;
Headings (most frequently used words):
and, to, kubernetes, kubeflow, istio, proxy, the, from, spark, recent, kemu, declarative, approach, emulating, clusters, at, scale, secure, ingress, authentication, with, external, auth, dex, oauth2, ultimate, homelab, guide, zero, production, cluster, on, premises, training, operators, solving, sidecar, lifecycle, problem, for, ai, ml, workloads, jobserver, standalone, mesos, marathon, docker,
Text of the page (most frequently used words):
and (26), the (17), kubernetes (12), for (11), with (9), #kubeflow (8), spark (7), istio (6), #jobserver (5), workloads (5), this (5), mins (5), multi (4), cluster (4), clusters (4), datastrophic (3), running (3), blog (3), post (3), software (3), more (3), 2021 (3), production (3), guide (3), configuration (3), publicly (3), exposed (3), kemu (3), about (3), 2025 (2), several (2), availability (2), tenancy (2), across (2), covers (2), made (2), installations (2), these (2), systems (2), docker (2), marathon (2), mesos (2), october (2), from (2), gaining (2), that (2), data (2), can (2), service (2), solving (2), provides (2), operators (2), proxy (2), end (2), part (2), provisioning (2), infrastructure (2), december (2), malicious (2), your (2), dashboard (2), workload (2), scheduling (2), secure (2), ingress (2), authentication (2), experimentation (2), declarative (2), scale (2), archive (2), after, years, need, better, emerged, projects, author, was, involved, design, decisions, provide, higher, fault, tolerance, scalability, failure, recovery, automation, choices, order, reach, goals, widely, used, variety, reporting, aggregating, 2017, standalone, traction, community, its, early, adoption, enterprises, security, observability, concerns, become, important, many, organizations, are, operate, sensitive, personal, financial, have, stricter, requirements, encryption, traceability, access, control, quite, often, see, use, mesh, problems, other, benefits, rich, functionality, mlops, training, sidecar, lifecycle, problem, whether, you, looking, powerful, development, environment, grade, experiments, deployment, instructions, get, first, planning, proxmox, terraform, second, dedicated, installing, essential, such, calico, networking, openebs, volume, metallb, network, load, balancing, devops, ultimate, homelab, zero, premises, insecure, endpoints, produce, major, risk, being, deployed, seen, reports, central, pipelines, all, were, compromised, when, internet, combined, wide, rbac, permissions, capabilities, opens, deployments, anybody, knowing, endpoint, url, focuses, building, stack, targeting, external, auth, dex, oauth2, optimizing, requires, extensive, observation, but, testing, scheduler, modifications, risky, errors, cause, day, delays, wasted, capacity, introduces, emulator, utility, replaces, fragmented, tool, setups, single, specification, enabling, safe, large, gpu, minimal, resources, kwok, kind, emulation, november, approach, emulating, recent, distributed, container, orchestration, platforms, skip, main, content,
Text of the page (random words):
datastrophic skip to main content datastrophic archive about archive about a blog about distributed systems container orchestration platforms and ai infrastructure recent kemu a declarative approach to emulating kubernetes clusters at scale 4 november 2025 12 mins kubernetes emulation kind kwok kemu optimizing ai workload scheduling requires extensive experimentation and observation but testing scheduler modifications in production is risky configuration errors can cause multi day delays and wasted capacity this post introduces kemu a declarative kubernetes emulator utility that replaces fragmented multi tool cluster setups with a single configuration specification enabling safe experimentation with large scale gpu clusters on minimal resources secure kubeflow ingress and authentication with istio external auth dex and oauth2 proxy 16 december 2021 16 mins kubernetes istio kubeflow publicly exposed insecure service endpoints on kubernetes produce a major risk of malicious workloads being deployed on your clusters we ve seen reports of the kubernetes dashboard the kubeflow central dashboard and the kubeflow pipelines all were compromised when publicly exposed to the internet combined with wide rbac permissions a publicly exposed software with workload scheduling capabilities opens your clusters for malicious deployments to anybody knowing the endpoint url this blog post focuses on building a secure ingress and authentication stack on kubernetes with istio targeting kubeflow installations the ultimate kubernetes homelab guide from zero to production cluster on premises 1 december 2021 14 mins kubernetes devops whether you re looking for a more powerful development environment or a production grade kubernetes cluster for experiments this guide provides end to end deployment and configuration instructions to get the cluster up and running the first part of this guide covers the planning and provisioning of the infrastructure with proxmox and terraform the second part is dedicated to installing kubernetes and essential software such as calico for networking openebs for volume provisioning and metallb for network load balancing kubeflow training operators and istio solving the proxy sidecar lifecycle problem for ai ml workloads 4 october 2021 12 mins kubernetes kubeflow istio operators mlops with kubeflow gaining traction in the community and its early adoption in enterprises security and observability concerns become more and more important many organizations that are running ai ml workloads operate with sensitive personal or financial data and have stricter requirements for data encryption traceability and access control quite often we can see the use of the istio service mesh for solving these problems and gaining other benefits of the rich functionality it provides spark jobserver from spark standalone to mesos marathon and docker 12 october 2017 9 mins spark mesos marathon docker after several years of running spark jobserver workloads the need for better availability and multi tenancy emerged across several projects author was involved in this blog post covers design decisions made to provide higher availability and fault tolerance of jobserver installations multi tenancy for spark workloads scalability and failure recovery automation and software choices made in order to reach these goals spark jobserver spark jobserver is widely used across a variety of reporting and aggregating systems 2025 datastrophic
|