[{"data":1,"prerenderedAt":412},["ShallowReactive",2],{"story":3},{"name":4,"created_at":5,"published_at":6,"updated_at":7,"id":8,"uuid":9,"content":10,"slug":403,"full_slug":404,"sort_by_date":25,"position":405,"tag_list":406,"is_startpage":241,"parent_id":407,"meta_data":25,"group_id":408,"first_published_at":409,"release_id":25,"lang":410,"path":25,"alternates":411,"default_full_slug":25,"translated_slugs":25},"AI-INFRA 100: AI Infrastructure","2026-09-18T10:58:13.501Z","2026-09-23T11:52:11.927Z","2026-09-23T11:52:11.942Z",221345765403475,"9f452448-d0ca-45fe-918a-3d3246df1250",{"seo":11,"_uid":15,"type":16,"intro":13,"title":17,"duration":18,"component":19,"course_id":20,"technology":21,"description":22,"on_schedule":241,"course_level":242,"hide_sidebar":241,"prerequisites":243,"on_demand_link":290,"lab_requirements":293,"course_objectives":326,"follow_up_courses":371,"who_should_attend":374,"additional_content":392,"on_demand_training":241,"section_below_hero":397,"hide_class_schedule_cta":402,"hide_private_training_cta":241,"section_below_hero_bg_color":13},{"_uid":12,"title":13,"plugin":14,"og_image":13,"og_title":13,"description":13,"twitter_image":13,"twitter_title":13,"og_description":13,"twitter_description":13},"867ef41d-665c-4afb-ab8c-7929189d5cb2","","seo_metatags","8fdb6453-e42e-443a-befd-e555c41d99b5","course","AI Infrastructure","2 days","training_course","AI-INFRA 100:","cn",{"type":23,"attrs":24,"content":26},"doc",{"backgroundColor":25},null,[27,34,39,48,54,93,98,142,147,183,188,231,236],{"type":28,"attrs":29,"content":30},"paragraph",{"textAlign":25},[31],{"text":32,"type":33},"A GPU cluster built for AI workloads runs on infrastructure that does not behave like a conventional datacenter. This introductory-level 2-day course covers that stack across four modules: fundamental architecture, fleet hardware management, network fabrics, and bare-metal GPU provisioning from Kubernetes. The goal is a working mental model of the whole stack, and hands-on exposure to the key components.","text",{"type":28,"attrs":35,"content":36},{"textAlign":25},[37],{"text":38,"type":33},"The four modules run in sequence, each building on the vocabulary and mental model established by the one before it. Three of the four include hands-on labs, on virtual environments that stand in for a real cluster.",{"type":40,"attrs":41,"content":43},"heading",{"level":42,"textAlign":25},2,[44,46],{"type":45},"hard_break",{"text":47,"type":33},"Course Outline",{"type":40,"attrs":49,"content":51},{"level":50,"textAlign":25},3,[52],{"text":53,"type":33},"Module 1: AI Infrastructure Foundations",{"type":55,"content":56},"bullet_list",[57,65,72,79,86],{"type":58,"content":59},"list_item",[60],{"type":28,"attrs":61,"content":62},{"textAlign":25},[63],{"text":64,"type":33},"The GPU node: accelerators, HBM, and the checkpoint burst",{"type":58,"content":66},[67],{"type":28,"attrs":68,"content":69},{"textAlign":25},[70],{"text":71,"type":33},"NVLink and intra-node interconnect",{"type":58,"content":73},[74],{"type":28,"attrs":75,"content":76},{"textAlign":25},[77],{"text":78,"type":33},"The four fabrics and what each one carries",{"type":58,"content":80},[81],{"type":28,"attrs":82,"content":83},{"textAlign":25},[84],{"text":85,"type":33},"The management plane",{"type":58,"content":87},[88],{"type":28,"attrs":89,"content":90},{"textAlign":25},[91],{"text":92,"type":33},"Failure domains at each layer",{"type":40,"attrs":94,"content":95},{"level":50,"textAlign":25},[96],{"text":97,"type":33},"Module 2: Redfish and Fleet Hardware Management",{"type":55,"content":99},[100,107,114,121,128,135],{"type":58,"content":101},[102],{"type":28,"attrs":103,"content":104},{"textAlign":25},[105],{"text":106,"type":33},"The Redfish resource model and mixed-vendor BMCs",{"type":58,"content":108},[109],{"type":28,"attrs":110,"content":111},{"textAlign":25},[112],{"text":113,"type":33},"Health validation before provisioning",{"type":58,"content":115},[116],{"type":28,"attrs":117,"content":118},{"textAlign":25},[119],{"text":120,"type":33},"Firmware compliance and drift",{"type":58,"content":122},[123],{"type":28,"attrs":124,"content":125},{"textAlign":25},[126],{"text":127,"type":33},"The scrape pipeline behind a fleet dashboard",{"type":58,"content":129},[130],{"type":28,"attrs":131,"content":132},{"textAlign":25},[133],{"text":134,"type":33},"VirtualMedia provisioning boot",{"type":58,"content":136},[137],{"type":28,"attrs":138,"content":139},{"textAlign":25},[140],{"text":141,"type":33},"Hands-on labs",{"type":40,"attrs":143,"content":144},{"level":50,"textAlign":25},[145],{"text":146,"type":33},"Module 3: GPU Cluster Networking: RoCE v2, PFC/ECN",{"type":55,"content":148},[149,156,163,170,177],{"type":58,"content":150},[151],{"type":28,"attrs":152,"content":153},{"textAlign":25},[154],{"text":155,"type":33},"Why RDMA cannot tolerate packet loss",{"type":58,"content":157},[158],{"type":28,"attrs":159,"content":160},{"textAlign":25},[161],{"text":162,"type":33},"PFC and ECN on the switch and on the NIC",{"type":58,"content":164},[165],{"type":28,"attrs":166,"content":167},{"textAlign":25},[168],{"text":169,"type":33},"Fabric telemetry and what it surfaces",{"type":58,"content":171},[172],{"type":28,"attrs":173,"content":174},{"textAlign":25},[175],{"text":176,"type":33},"Worked incidents traced from symptom to root cause",{"type":58,"content":178},[179],{"type":28,"attrs":180,"content":181},{"textAlign":25},[182],{"text":141,"type":33},{"type":40,"attrs":184,"content":185},{"level":50,"textAlign":25},[186],{"text":187,"type":33},"Module 4: Bare-Metal GPU Provisioning with Kubernetes",{"type":55,"content":189},[190,197,204,211,218,225],{"type":58,"content":191},[192],{"type":28,"attrs":193,"content":194},{"textAlign":25},[195],{"text":196,"type":33},"Host enrollment and hardware inventory",{"type":58,"content":198},[199],{"type":28,"attrs":200,"content":201},{"textAlign":25},[202],{"text":203,"type":33},"Declarative provisioning with k0rdent and Metal3/Ironic",{"type":58,"content":205},[206],{"type":28,"attrs":207,"content":208},{"textAlign":25},[209],{"text":210,"type":33},"GPU Operator and Network Operator on a mixed fleet",{"type":58,"content":212},[213],{"type":28,"attrs":214,"content":215},{"textAlign":25},[216],{"text":217,"type":33},"Incremental deploy-and-validate methodology",{"type":58,"content":219},[220],{"type":28,"attrs":221,"content":222},{"textAlign":25},[223],{"text":224,"type":33},"Tracing a provisioning failure to its hardware root cause",{"type":58,"content":226},[227],{"type":28,"attrs":228,"content":229},{"textAlign":25},[230],{"text":141,"type":33},{"type":40,"attrs":232,"content":233},{"level":42,"textAlign":25},[234],{"text":235,"type":33},"Format",{"type":28,"attrs":237,"content":238},{"textAlign":25},[239],{"text":240,"type":33},"Instructor-led, two consecutive days, four sequential modules. Hands-on labs in modules 2, 3, and 4 on hosted, preconfigured environments. Reference materials are yours to keep. Available as a private delivery for your team.",false,"essentials",{"type":23,"attrs":244,"content":245},{"backgroundColor":25},[246],{"type":55,"content":247},[248,255,262,269,276,283],{"type":58,"content":249},[250],{"type":28,"attrs":251,"content":252},{"textAlign":25},[253],{"text":254,"type":33},"Solid Linux command line proficiency",{"type":58,"content":256},[257],{"type":28,"attrs":258,"content":259},{"textAlign":25},[260],{"text":261,"type":33},"TCP/IP networking, VLANs, and L2/L3 switching fundamentals",{"type":58,"content":263},[264],{"type":28,"attrs":265,"content":266},{"textAlign":25},[267],{"text":268,"type":33},"Server hardware fundamentals, including exposure to BMC or IPMI",{"type":58,"content":270},[271],{"type":28,"attrs":272,"content":273},{"textAlign":25},[274],{"text":275,"type":33},"Ability to read JSON and YAML, and use curl",{"type":58,"content":277},[278],{"type":28,"attrs":279,"content":280},{"textAlign":25},[281],{"text":282,"type":33},"Kubernetes and kubectl basics, for the provisioning module",{"type":58,"content":284},[285],{"type":28,"attrs":286,"content":287},{"textAlign":25},[288],{"text":289,"type":33},"Beneficial (not mandatory): shell scripting, and familiarity with Kubernetes operators",{"id":13,"url":13,"linktype":291,"fieldtype":292,"cached_url":13},"story","multilink",{"type":23,"attrs":294,"content":295},{"backgroundColor":25},[296],{"type":55,"content":297},[298,305,312,319],{"type":58,"content":299},[300],{"type":28,"attrs":301,"content":302},{"textAlign":25},[303],{"text":304,"type":33},"WiFi-enabled laptop",{"type":58,"content":306},[307],{"type":28,"attrs":308,"content":309},{"textAlign":25},[310],{"text":311,"type":33},"Current Chrome or Firefox browser",{"type":58,"content":313},[314],{"type":28,"attrs":315,"content":316},{"textAlign":25},[317],{"text":318,"type":33},"SSH client",{"type":58,"content":320},[321],{"type":28,"attrs":322,"content":323},{"textAlign":25},[324],{"text":325,"type":33},"Lab environments are hosted and preconfigured; no local installation or hardware is required",{"type":23,"attrs":327,"content":328},{"backgroundColor":25},[329,334,344,353,362],{"type":28,"attrs":330,"content":331},{"textAlign":25},[332],{"text":333,"type":33},"The curriculum encompasses four module domains:",{"type":28,"attrs":335,"content":336},{"textAlign":25},[337,342],{"text":338,"type":33,"marks":339},"AI Infrastructure Foundations:",[340],{"type":341},"bold",{"text":343,"type":33}," A conceptual walkthrough of the GPU cluster stack, covering the GPU node, HBM and the checkpoint burst, NVLink and intra-node interconnect, the four fabrics and what each one carries, the management plane, and the failure domains at each layer. Introduces the vocabulary and architecture the remaining three modules assume, and gives participants enough grounding to follow a vendor conversation without nodding at words they cannot define. Conceptual; no lab.",{"type":28,"attrs":345,"content":346},{"textAlign":25},[347,351],{"text":348,"type":33,"marks":349},"Redfish and Fleet Hardware Management:",[350],{"type":341},{"text":352,"type":33}," An introduction to hardware management and observability through one HTTPS API across a mixed-vendor fleet. Covers the Redfish resource model, health validation before provisioning, firmware compliance and drift, the scrape pipeline behind a fleet dashboard, and VirtualMedia provisioning boot. The multi-vendor problem is the practical heart of the module: every BMC vendor has its own interface and its own quirks, and clicking through four different web UIs does not survive contact with a real fleet. Hands-on labs.",{"type":28,"attrs":354,"content":355},{"textAlign":25},[356,360],{"text":357,"type":33,"marks":358},"GPU Cluster Networking: RoCE v2, PFC/ECN:",[359],{"type":341},{"text":361,"type":33}," An introduction to lossless fabric, covering why RDMA cannot tolerate loss, what PFC and ECN do on both the switch and the NIC, fabric telemetry, and worked incidents traced from symptom to root cause. Participants leave able to reason about fabric behavior and read fabric telemetry, not to design a production fabric unaided. Hands-on labs.",{"type":28,"attrs":363,"content":364},{"textAlign":25},[365,369],{"text":366,"type":33,"marks":367},"Bare-Metal GPU Provisioning with Kubernetes:",[368],{"type":341},{"text":370,"type":33}," An introduction to declarative bare-metal provisioning. Covers host enrollment, provisioning with k0rdent and Metal3/Ironic, GPU Operator and Network Operator deployment on a mixed fleet, an incremental deploy-and-validate methodology, and a provisioning failure traced back to its hardware root cause. Hands-on labs.",{"type":23,"content":372},[373],{"type":28},{"type":23,"attrs":375,"content":376},{"backgroundColor":25},[377,382,387],{"type":28,"attrs":378,"content":379},{"textAlign":25},[380],{"text":381,"type":33},"The course addresses infrastructure engineers, datacenter and platform operations staff, and field, solutions, and pre-sales engineers.",{"type":28,"attrs":383,"content":384},{"textAlign":25},[385],{"text":386,"type":33},"Infrastructure engineers encountering GPU cluster work for the first time get the map and the vocabulary before the first outage. Platform and operations teams, inheriting a cluster somebody else stood up, get a model of what they have been handed. Field and pre-sales roles who need to follow a technical conversation about GPU topology, RDMA fabric, and cluster automation without implementing it get most of that value from the first two modules.",{"type":28,"attrs":388,"content":389},{"textAlign":25},[390],{"text":391,"type":33},"Participants should possess practical datacenter experience with servers, switching, and BMC-based hardware management. No RDMA, DPU, Redfish, or GPU cluster experience is assumed.",{"type":23,"attrs":393,"content":394},{"backgroundColor":25},[395],{"type":28,"attrs":396},{"textAlign":25},{"type":23,"attrs":398,"content":399},{"backgroundColor":25},[400],{"type":28,"attrs":401},{"textAlign":25},true,"ai-infrastructure","training/courses/ai-infrastructure",-300,[],111882102,"5752396c-3905-4d27-96b3-f3162d73c931","2026-09-18T11:03:23.769Z","default",[],1790594830444]