{"id":76120,"date":"2025-12-04T04:37:38","date_gmt":"2025-12-03T23:07:38","guid":{"rendered":"https:\/\/www.tothenew.com\/blog\/?p=76120"},"modified":"2026-09-15T16:19:55","modified_gmt":"2026-09-15T10:49:55","slug":"building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades","status":"publish","type":"post","link":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/","title":{"rendered":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades"},"content":{"rendered":"<p><strong>Introduction<\/strong><\/p>\n<p>Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate.<\/p>\n<p>This is a working reference: routines you&#8217;ll actually run, what a rolling upgrade really does, the concepts worth internalizing, and the metrics that separate &#8220;the cluster is up&#8221; from genuinely competent operation.<\/p>\n<p>In this blog, you can expect to read about pragmatic methods for achieving competency in Kafka:<\/p>\n<ul>\n<li>Daily Operations with Kafka<\/li>\n<li>Upgrading MSK<\/li>\n<li>Deepen your Understanding of Kafka Partitions, Topics, Producers, Consumers<\/li>\n<li>Monitoring Key Metrics<\/li>\n<\/ul>\n<p><strong>1. Routine Kafka Operations<\/strong><\/p>\n<p>Running Kafka clusters involves a blend of everyday responsibilities and incident response-associated readiness. Some of the main categories include:<\/p>\n<p><strong>a) Topic Administration<\/strong><\/p>\n<p>Partition count and replication factor are far cheaper to get right on day one \u2014 partitions can be added but never removed. Settle on a naming convention (appname-env-eventtype) before the tenth team invents its own.<\/p>\n<blockquote><p>kafka-topics.sh &#8211;bootstrap-server $BROKERS &#8211;command-config client.properties &#8211;create &#8211;topic user-events &#8211;partitions 6 &#8211;replication-factor 3<\/p><\/blockquote>\n<ul>\n<li><code><\/code>Size for the consumer you&#8217;ll have \u2014 partition count caps read parallelism.<\/li>\n<li>Replicate across every AZ the cluster spans \u2014 RF 2 on a three-AZ cluster still risks losing both copies to one AZ event.<\/li>\n<li>Grow partitions, never shrink them\u00a0\u2014 increasing is live; decreasing isn&#8217;t supported.<\/li>\n<\/ul>\n<p><strong>b) Consumer Group Monitoring<\/strong><\/p>\n<p>Consumer lag is the single most useful number in Kafka ops \u2014 a consumer falling behind before anyone notices downstream. CloudWatch covers the alarm;\u00a0kafka-consumer-groups.sh\u00a0shows which partitions are behind.<\/p>\n<p><code>kafka-consumer-groups.sh --bootstrap-server &lt;broker&gt;<br \/>\n--describe --group user-event-processor<\/code><\/p>\n<ul>\n<li>An even partition-to-member ratio \u2014 one member starved is usually a hot key, not capacity.<\/li>\n<li>Lag that climbs steadily vs. spikes and recovers \u2014 &#8220;too slow&#8221; vs. &#8220;a blip.&#8221;<\/li>\n<li>Rebalance frequency\u00a0\u2014 rebalancing every few minutes is a real availability tax.<\/li>\n<\/ul>\n<p><strong>c) ACLs and Security<\/strong><\/p>\n<p>MSK supports both SASL\/SCRAM and IAM access control, and they solve slightly different problems. IAM access control maps naturally onto infrastructure that already lives in IAM \u2014 a Lambda function or an ECS task role gets Kafka access the same way it gets S3 access, with no separate credential to rotate. SASL\/SCRAM earns its keep for external or non-AWS clients that have no IAM identity to assume. Either way, the ACL itself should be scoped to an operation, not a blanket grant:<\/p>\n<ul>\n<li>READ \u2014 consume from a specific topic or group, never cluster-wide.<\/li>\n<li>WRITE \u2014 produce to a named topic; pair with a quota so one misbehaving producer can&#8217;t saturate broker throughput for everyone else.<\/li>\n<li>DESCRIBE\u00a0\u2014 metadata visibility, commonly granted more liberally since it exposes shape, not data.<\/li>\n<\/ul>\n<p><strong>d) Retention and Archival<\/strong><\/p>\n<p>Per-topic retention should reflect how the data is actually used, not a single cluster-wide default \u2014 a clickstream topic feeding a real-time model might only need 24 hours, while an audit-relevant event topic might need 30 days or more. Two mechanisms extend that beyond what&#8217;s practical to keep on broker disk: MSK&#8217;s Tiered Storage moves older segments to a low-cost tier while they&#8217;re still addressable through the normal Kafka consumer API, and a Kafka Connect S3 sink instead exports data out of Kafka entirely for long-term, queryable archival. They&#8217;re not competing options \u2014 Tiered Storage keeps Kafka the interface, the S3 sink hands the data to something else.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-83223 size-full\" src=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram.png\" alt=\"topic-anatomy-diagram\" width=\"1440\" height=\"782\" srcset=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram.png 1440w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram-300x163.png 300w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram-1024x556.png 1024w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram-768x417.png 768w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram-624x339.png 624w\" sizes=\"auto, (max-width: 1440px) 100vw, 1440px\" \/><\/p>\n<p><strong>2. MSK Upgrades<\/strong><\/p>\n<p>Amazon MSK simplifies upgrades by providing broker patching and managing rolling upgrades so that engineers can make intelligent decisions.<\/p>\n<p><strong>a) Plan before you roll<\/strong><\/p>\n<p>Treat every upgrade as a compatibility check first, an operational event second. Kafka guarantees client\/broker compatibility in both directions across supported versions, but that guarantee doesn&#8217;t extend to your own tooling \u2014 Kafka Connect connectors and Kafka Streams applications have their own version constraints, and a deprecated config silently ignored is worse than one that fails loudly.<\/p>\n<ul>\n<li>Run the target version in a QA or staging cluster first \u2014 long enough to catch anything version-specific in your own producer\/consumer code, not just a smoke test.<\/li>\n<li>Read the release notes for breaking protocol changes, deprecated configs, and default-value changes \u2014 not just the headline features.<\/li>\n<li>Check Kafka Connect and Kafka Streams compatibility\u00a0against the target version before scheduling anything in production.<\/li>\n<\/ul>\n<p><strong>b) What actually happens during the rolling upgrade<\/strong><\/p>\n<p>MSK upgrades a cluster broker by broker, never all at once. Each broker briefly leaves service for its own patch window while the rest of the cluster keeps serving reads and writes \u2014 there&#8217;s no cluster-wide downtime, but there is a moment where one broker&#8217;s partitions are down to their remaining in-sync replicas. Two settings decide whether your clients notice: replication factor and\u00a0min.insync.replicas. On a three-AZ cluster, run\u00a0RF 3 with\u00a0min.insync.replicas\u00a0set to 2\u00a0\u2014 that tolerates exactly one broker being mid-patch without blocking writes. RF 2 doesn&#8217;t have that slack; losing one replica during patching can push a partition to zero in-sync replicas and take writes offline for it.<\/p>\n<blockquote><p>Where this actually bitesThe failure mode isn&#8217;t the upgrade itself \u2014 it&#8217;s a topic someone created months ago with the default replication factor, on a cluster that&#8217;s since grown to three AZs. Audit RF and minISR before you schedule the upgrade, not after a partition goes offline mid-patch.<\/p><\/blockquote>\n<p>On the client side, configure every producer and consumer with the full list of broker addresses, not just one. A client that only knows about the broker currently being patched has nothing to fail over to; a client that knows the whole cluster reconnects to a different broker automatically and keeps going.<\/p>\n<p><strong>3. After the upgrade: rebalancing<\/strong><\/p>\n<p>Whether you scaled the cluster or just rolled the version, partitions can end up unevenly distributed across brokers. How you fix that depends on the broker type:\u00a0Express brokers\u00a0handle this automatically through Intelligent Rebalancing, which redistributes partitions after a scaling event with no configuration and no third-party tooling.\u00a0Standard brokers\u00a0need it done explicitly, either with the open-source\u00a0kafka-reassign-partitions.sh\u00a0tool or with Cruise Control, which automates the same reassignment based on live load rather than a manual plan.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-83222 size-full\" src=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram.png\" alt=\"rolling-upgrade-diagram\" width=\"1440\" height=\"642\" srcset=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram.png 1440w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram-300x134.png 300w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram-1024x457.png 1024w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram-768x342.png 768w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/rolling-upgrade-diagram-624x278.png 624w\" sizes=\"auto, (max-width: 1440px) 100vw, 1440px\" \/><\/p>\n<p><strong>4. The four concepts everything else rests on<\/strong><\/p>\n<ul>\n<li>Topics are named logical streams. No ordering guarantee on their own \u2014 ordering exists only within a partition.<\/li>\n<li>Partitions are the unit of parallelism and ordering \u2014 an append-only log, strictly ordered within itself, unordered across partitions.<\/li>\n<li>Producers pick the partition a record lands in, by key hash (default \u2014 gives per-key ordering), round-robin, or a custom partitioner.<\/li>\n<li>Consumers\u00a0read partitions in groups; a partition is read by at most one member at a time.<\/li>\n<\/ul>\n<p><strong>b) Worked Example<\/strong><\/p>\n<p>Topic\u00a0user-activity, six partitions, keyed by user ID \u2014 every event for a user lands in order in the same partition.\u00a0fraud-detect\u00a0reads it at three members (two partitions each);\u00a0analytics\u00a0at six (one each, for throughput).<\/p>\n<p><strong>c) Sizing partitions in practice<\/strong><\/p>\n<p>Too few caps parallelism for every group that will read the topic. Too many costs file handles, replication traffic, and slower leader elections on failure. Start at 2\u20133x expected peak parallelism, then grow \u2014 never shrink.<br \/>\n<strong>5. Watching the right metrics<\/strong><\/p>\n<p>A cluster that looks healthy in the console can still be quietly failing one consumer or one partition. These metrics catch that gap.<\/p>\n<ul>\n<li>Broker-level metrics\n<ul>\n<li>BytesInPerSec \/ BytesOutPerSec \u2014 throughput per broker; a sustained skew usually means uneven leader distribution.<\/li>\n<li>UnderReplicatedPartitions \u2014 zero outside a patch window; persistent nonzero is worth paging on.<\/li>\n<li>ActiveControllerCount\u00a0\u2014 always exactly 1. Zero or more than 1 both signal trouble.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<ul>\n<li>Topic and partition metrics\n<ul>\n<li>Consumer lag \u2014 a partition&#8217;s latest offset minus a group&#8217;s committed one; the best leading indicator of trouble.<\/li>\n<li>PartitionCount \u2014 tracked over time for capacity planning.<\/li>\n<li>OfflinePartitionsCount\u00a0\u2014 always zero, or a partition has no leader.<\/li>\n<\/ul>\n<\/li>\n<\/ul>\n<p><strong>6. Choosing a CloudWatch monitoring level<\/strong><\/p>\n<p>MSK&#8217;s CloudWatch monitoring is tiered \u2014\u00a0DEFAULT\u00a0covers cluster and broker metrics;\u00a0PER_TOPIC_PER_BROKER\u00a0and\u00a0PER_TOPIC_PER_PARTITION\u00a0go finer, at real added cost. Reserve those for topics you alarm on individually, and let automation handle the routine responses \u2014 a lag alarm triggering a Lambda beats a page for something a script could fix.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-83221 size-full\" src=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup.png\" alt=\"monitoring-dashboard-mockup\" width=\"1440\" height=\"542\" srcset=\"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup.png 1440w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup-300x113.png 300w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup-1024x385.png 1024w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup-768x289.png 768w, https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/monitoring-dashboard-mockup-624x235.png 624w\" sizes=\"auto, (max-width: 1440px) 100vw, 1440px\" \/><\/p>\n<p><strong>7. Building real competence<\/strong><\/p>\n<ul>\n<li>Run fire drills. Kill a broker in staging and watch leadership, ISR, and clients react \u2014 the first rebalance you see shouldn&#8217;t be a real incident.<\/li>\n<li>Manage topics and ACLs as code, so the config that&#8217;s running is the config in version control.<\/li>\n<li>Let alarms trigger real remediation \u2014 a Lambda reacting to a lag alarm resolves pages before a human sees them.<\/li>\n<li>Rehearse every upgrade in QA first, where a Connect plugin&#8217;s incompatibility should surface before production.<\/li>\n<\/ul>\n<p><strong>Conclusion<\/strong><\/p>\n<p>Running Kafka well on MSK comes down to a few things: size topics right the first time, understand what a rolling upgrade does to replication, watch metrics that reveal a localized problem instead of a green light, and practice failure before production forces you to learn it live.<\/p>\n<ul>\n<li>Get naming, partition count, and RF right before the first producer writes to a topic.<\/li>\n<li>min.insync.replicas and full broker-address client configs are upgrade prerequisites.<\/li>\n<li>Alarm on consumer lag and under-replication \u2014 a healthy-looking cluster can still be failing one group.<\/li>\n<li>Rehearse failure in staging so the first broker outage you see isn&#8217;t in production.<\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you&#8217;ll actually run, what a rolling upgrade really does, the concepts worth internalizing, [&hellip;]<\/p>\n","protected":false},"author":1719,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"iawp_total_views":2,"footnotes":""},"categories":[2348],"tags":[1892,1604],"class_list":["post-76120","post","type-post","status-publish","format-standard","hentry","category-devops-technology","tag-devops","tag-kafka"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.0.1 - aioseo.com -->\n\t<meta name=\"description\" content=\"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you&#039;ll actually run, what a rolling upgrade really does, the concepts worth internalizing,\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Saif Ahmad\"\/>\n\t<link rel=\"canonical\" href=\"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.0.1\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"TO THE NEW BLOG\" \/>\n\t\t<meta property=\"og:type\" content=\"blog\" \/>\n\t\t<meta property=\"og:title\" content=\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog\" \/>\n\t\t<meta property=\"og:description\" content=\"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you&#039;ll actually run, what a rolling upgrade really does, the concepts worth internalizing,\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/\" \/>\n\t\t<meta property=\"og:image\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta property=\"og:image:secure_url\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary\" \/>\n\t\t<meta name=\"twitter:site\" content=\"@tothenew\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog\" \/>\n\t\t<meta name=\"twitter:description\" content=\"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you&#039;ll actually run, what a rolling upgrade really does, the concepts worth internalizing,\" \/>\n\t\t<meta name=\"twitter:image\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#article\",\"name\":\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog\",\"headline\":\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades\",\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/saif-ahmad\\\/#author\"},\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/wp-ttn-blog\\\/uploads\\\/2026\\\/09\\\/topic-anatomy-diagram.png\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#articleImage\"},\"datePublished\":\"2025-12-04T04:37:38+05:30\",\"dateModified\":\"2026-09-15T16:19:55+05:30\",\"inLanguage\":\"en-US\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#webpage\"},\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#webpage\"},\"articleSection\":\"DevOps, devops, Kafka\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/devops-technology\\\/#listItem\",\"name\":\"DevOps\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/devops-technology\\\/#listItem\",\"position\":2,\"name\":\"DevOps\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/devops-technology\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#listItem\",\"name\":\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#listItem\",\"position\":3,\"name\":\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/devops-technology\\\/#listItem\",\"name\":\"DevOps\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\",\"name\":\"TO THE NEW Blog\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/saif-ahmad\\\/#author\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/saif-ahmad\\\/\",\"name\":\"Saif Ahmad\",\"image\":{\"@type\":\"ImageObject\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#authorImage\",\"url\":\"https:\\\/\\\/newersworld-sf-static.tothenew.net\\\/prod\\\/profilePicFolder\\\/6d7b39d4-2956-4c96-bbfc-d3d04b4d31a1_5084-Saif-Ahmad-PROFILEPICTURE.jpeg\",\"width\":96,\"height\":96,\"caption\":\"Saif Ahmad\"}},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#webpage\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/\",\"name\":\"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog\",\"description\":\"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you'll actually run, what a rolling upgrade really does, the concepts worth internalizing,\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\\\/#breadcrumblist\"},\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/saif-ahmad\\\/#author\"},\"creator\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/saif-ahmad\\\/#author\"},\"datePublished\":\"2025-12-04T04:37:38+05:30\",\"dateModified\":\"2026-09-15T16:19:55+05:30\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\",\"name\":\"TO THE NEW Blog\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog","description":"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you'll actually run, what a rolling upgrade really does, the concepts worth internalizing,","canonical_url":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#article","name":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog","headline":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades","author":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/saif-ahmad\/#author"},"publisher":{"@id":"https:\/\/www.tothenew.com\/blog\/#organization"},"image":{"@type":"ImageObject","url":"https:\/\/www.tothenew.com\/blog\/wp-ttn-blog\/uploads\/2026\/09\/topic-anatomy-diagram.png","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#articleImage"},"datePublished":"2025-12-04T04:37:38+05:30","dateModified":"2026-09-15T16:19:55+05:30","inLanguage":"en-US","mainEntityOfPage":{"@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#webpage"},"isPartOf":{"@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#webpage"},"articleSection":"DevOps, devops, Kafka"},{"@type":"BreadcrumbList","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog#listItem","position":1,"name":"Home","item":"https:\/\/www.tothenew.com\/blog","nextItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/#listItem","name":"DevOps"}},{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/#listItem","position":2,"name":"DevOps","item":"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/","nextItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#listItem","name":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades"},"previousItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#listItem","position":3,"name":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades","previousItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/#listItem","name":"DevOps"}}]},{"@type":"Organization","@id":"https:\/\/www.tothenew.com\/blog\/#organization","name":"TO THE NEW Blog","url":"https:\/\/www.tothenew.com\/blog\/"},{"@type":"Person","@id":"https:\/\/www.tothenew.com\/blog\/author\/saif-ahmad\/#author","url":"https:\/\/www.tothenew.com\/blog\/author\/saif-ahmad\/","name":"Saif Ahmad","image":{"@type":"ImageObject","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#authorImage","url":"https:\/\/newersworld-sf-static.tothenew.net\/prod\/profilePicFolder\/6d7b39d4-2956-4c96-bbfc-d3d04b4d31a1_5084-Saif-Ahmad-PROFILEPICTURE.jpeg","width":96,"height":96,"caption":"Saif Ahmad"}},{"@type":"WebPage","@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#webpage","url":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/","name":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog","description":"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you'll actually run, what a rolling upgrade really does, the concepts worth internalizing,","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/www.tothenew.com\/blog\/#website"},"breadcrumb":{"@id":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/#breadcrumblist"},"author":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/saif-ahmad\/#author"},"creator":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/saif-ahmad\/#author"},"datePublished":"2025-12-04T04:37:38+05:30","dateModified":"2026-09-15T16:19:55+05:30"},{"@type":"WebSite","@id":"https:\/\/www.tothenew.com\/blog\/#website","url":"https:\/\/www.tothenew.com\/blog\/","name":"TO THE NEW Blog","inLanguage":"en-US","publisher":{"@id":"https:\/\/www.tothenew.com\/blog\/#organization"}}]},"og:locale":"en_US","og:site_name":"TO THE NEW BLOG","og:type":"blog","og:title":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog","og:description":"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you'll actually run, what a rolling upgrade really does, the concepts worth internalizing,","og:url":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/","og:image":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","og:image:secure_url":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","twitter:card":"summary","twitter:site":"@tothenew","twitter:title":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades | TO THE NEW Blog","twitter:description":"Introduction Apache Kafka is the default substrate for event-driven systems. Amazon MSK takes the operational floor out from under most of it (provisioning, patching, storage), but partitions, consumer groups, and replication are still yours to operate. This is a working reference: routines you'll actually run, what a rolling upgrade really does, the concepts worth internalizing,","twitter:image":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png"},"aioseo_meta_data":{"post_id":"76120","title":null,"description":null,"keywords":null,"keyphrases":{"focus":{"keyphrase":"","score":0,"analysis":{"keyphraseInTitle":{"score":0,"maxScore":9,"error":1}}},"additional":[]},"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":"","og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"Article","isEnabled":true},"graphs":[]},"schema_type":"default","schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":"-1","robots_max_videopreview":"-1","robots_max_imagepreview":"large","priority":null,"frequency":"default","local_seo":null,"limit_modified_date":false,"created":"2025-09-09 08:39:03","updated":"2026-09-15 10:49:57","focus_keyword":null,"additional_keywords":null,"truseo_locale":null,"ai":{"faqs":[],"keyPoints":[],"schemas":[],"titles":[],"descriptions":[],"socialPosts":{"email":{"subject":"","preview":"","content":""},"linkedin":[],"twitter":[],"facebook":[],"instagram":[]}},"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.tothenew.com\/blog\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/\" title=\"DevOps\">DevOps<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tBuilding Competence in Kafka: From Day-to-Day Operations to MSK Upgrades\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/www.tothenew.com\/blog"},{"label":"DevOps","link":"https:\/\/www.tothenew.com\/blog\/category\/devops-technology\/"},{"label":"Building Competence in Kafka: From Day-to-Day Operations to MSK Upgrades","link":"https:\/\/www.tothenew.com\/blog\/building-competence-in-kafka-from-day-to-day-operations-to-msk-upgrades\/"}],"_links":{"self":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts\/76120","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/users\/1719"}],"replies":[{"embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/comments?post=76120"}],"version-history":[{"count":16,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts\/76120\/revisions"}],"predecessor-version":[{"id":83472,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts\/76120\/revisions\/83472"}],"wp:attachment":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/media?parent=76120"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/categories?post=76120"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/tags?post=76120"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}