{"id":5754,"date":"2012-06-14T17:42:24","date_gmt":"2012-06-14T12:12:24","guid":{"rendered":"http:\/\/www.tothenew.com\/blog\/?p=5754"},"modified":"2016-12-19T15:16:40","modified_gmt":"2016-12-19T09:46:40","slug":"normalizing-accented-words","status":"publish","type":"post","link":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/","title":{"rendered":"Normalizing Accented Words"},"content":{"rendered":"<p>We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &#8220;saint&#8221; will always be normalized to &#8220;st.&#8221; etc.<\/p>\n<p>An important part of such normalization is to account for &#8216;accented&#8217; characters like \u00e9 or \u00e8 , and generally you would want them to normalize to normal\u00a0English\u00a0alphabet &#8216;e&#8217;, as that would help in sorting\/searching words containing these characters. For eg : you would want &#8220;Indianap\u00f2lis&#8221; to be normalized to &#8220;Indianapolis&#8221;.<\/p>\n<p>We can achieve this by using <strong>java.text.Normalizer<\/strong> class, all we need to do is<\/p>\n<p>[java]<br \/>\nNormalizer.normalize(&amp;quot;Indianap\u00f2lis&amp;quot;, Normalizer.Form.NFKD).replaceAll(&amp;quot;\\\\p{InCombiningDiacriticalMarks}+&amp;quot;, &amp;quot;&amp;quot;)<br \/>\n[\/java]<\/p>\n<p>Lets understand what is happening here, clearly we are calling a static function normalize in <strong>java.text.Normalizer<\/strong> class, the first parameter we passed is the string we want to normalize, the second parameter is the <a href=\"http:\/\/www.tothenew.com\/blog\/normalization-forms-for-accented-characters-in-java\/\">normalization form<\/a>. There are 4 normalization forms<\/p>\n<ol>\n<li>NFC &#8211; Canonical Decomposition, followed by Canonical Composition.<\/li>\n<li>NFD &#8211; Canonical Decomposition<\/li>\n<li>NFKC &#8211; Compatibility Decomposition, followed by Canonical Composition<\/li>\n<li>NFKD &#8211; Compatibility Decomposition<\/li>\n<\/ol>\n<p>So in the second parameter we pass in the NFKD form, which is an enum of type\u00a0<strong>Form<\/strong>.<\/p>\n<p>The normalizer function will still return us &#8220;Indianap\u00f2lis&#8221;. so, what happened there?<br \/>\n<br \/>\nLets understand, we can create an \u00f2 using two ways. It can either be a unicode character (U+00F2) or it can be a normal english &#8216;o&#8217; with a grave accent added to it (o+ (U+0060)).<br \/>\n<br \/>What normalizer function did was it normalized both cases to english &#8216;o&#8217; with an accent (o+ (U+0060))<br \/>\n<br \/>The next part is a replaceAll with a regex, the\u00a0{InCombiningDiacriticalMarks} is a Unicode block property, which matches the accented characters.<br \/>\n<br \/>The second part replaces all accented characters (grave accent (U+0060) in this case) with empty string. So what we finally get is &#8220;Indianapolis&#8221;, and we are done.<br \/>\n<br \/>\nHope it helped!<\/p>\n<p>Sachin Anand<\/p>\n<p>sachin[at]intelligrape[dot]com<\/p>\n<p>@babasachinanand<\/p>\n","protected":false},"excerpt":{"rendered":"<p>We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &#8220;saint&#8221; will always [&hellip;]<\/p>\n","protected":false},"author":14,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"iawp_total_views":11,"footnotes":""},"categories":[1],"tags":[829,830,828,827,831],"class_list":["post-5754","post","type-post","status-publish","format-standard","hentry","category-technology","tag-nfc","tag-nfkc","tag-nfkd","tag-normalization","tag-text-normalization"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.0.1 - aioseo.com -->\n\t<meta name=\"description\" content=\"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &quot;saint&quot; will always\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Sachin\"\/>\n\t<link rel=\"canonical\" href=\"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.0.1\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"TO THE NEW BLOG\" \/>\n\t\t<meta property=\"og:type\" content=\"blog\" \/>\n\t\t<meta property=\"og:title\" content=\"Normalizing Accented Words | TO THE NEW Blog\" \/>\n\t\t<meta property=\"og:description\" content=\"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &quot;saint&quot; will always\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/\" \/>\n\t\t<meta property=\"og:image\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta property=\"og:image:secure_url\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary\" \/>\n\t\t<meta name=\"twitter:site\" content=\"@tothenew\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Normalizing Accented Words | TO THE NEW Blog\" \/>\n\t\t<meta name=\"twitter:description\" content=\"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &quot;saint&quot; will always\" \/>\n\t\t<meta name=\"twitter:image\" content=\"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#article\",\"name\":\"Normalizing Accented Words | TO THE NEW Blog\",\"headline\":\"Normalizing Accented Words\",\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/sachin\\\/#author\"},\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"},\"datePublished\":\"2012-06-14T17:42:24+05:30\",\"dateModified\":\"2016-12-19T15:16:40+05:30\",\"inLanguage\":\"en-US\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#webpage\"},\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#webpage\"},\"articleSection\":\"Technology, NFC, NFKC, NFKD, Normalization, Text Normalization\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/technology\\\/#listItem\",\"name\":\"Technology\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/technology\\\/#listItem\",\"position\":2,\"name\":\"Technology\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/technology\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#listItem\",\"name\":\"Normalizing Accented Words\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#listItem\",\"position\":3,\"name\":\"Normalizing Accented Words\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/technology\\\/#listItem\",\"name\":\"Technology\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\",\"name\":\"TO THE NEW Blog\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/sachin\\\/#author\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/sachin\\\/\",\"name\":\"Sachin\",\"image\":{\"@type\":\"ImageObject\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#authorImage\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/d84b0b7a42bd13ad65b8a665e117addb0455be6b706c443a53526dc9d0e5af87?s=96&d=mm&r=g\",\"width\":96,\"height\":96,\"caption\":\"Sachin\"}},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#webpage\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/\",\"name\":\"Normalizing Accented Words | TO THE NEW Blog\",\"description\":\"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like \\\"saint\\\" will always\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/normalizing-accented-words\\\/#breadcrumblist\"},\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/sachin\\\/#author\"},\"creator\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/sachin\\\/#author\"},\"datePublished\":\"2012-06-14T17:42:24+05:30\",\"dateModified\":\"2016-12-19T15:16:40+05:30\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\",\"name\":\"TO THE NEW Blog\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Normalizing Accented Words | TO THE NEW Blog","description":"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like \"saint\" will always","canonical_url":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#article","name":"Normalizing Accented Words | TO THE NEW Blog","headline":"Normalizing Accented Words","author":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/sachin\/#author"},"publisher":{"@id":"https:\/\/www.tothenew.com\/blog\/#organization"},"datePublished":"2012-06-14T17:42:24+05:30","dateModified":"2016-12-19T15:16:40+05:30","inLanguage":"en-US","mainEntityOfPage":{"@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#webpage"},"isPartOf":{"@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#webpage"},"articleSection":"Technology, NFC, NFKC, NFKD, Normalization, Text Normalization"},{"@type":"BreadcrumbList","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog#listItem","position":1,"name":"Home","item":"https:\/\/www.tothenew.com\/blog","nextItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/technology\/#listItem","name":"Technology"}},{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/technology\/#listItem","position":2,"name":"Technology","item":"https:\/\/www.tothenew.com\/blog\/category\/technology\/","nextItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#listItem","name":"Normalizing Accented Words"},"previousItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#listItem","position":3,"name":"Normalizing Accented Words","previousItem":{"@type":"ListItem","@id":"https:\/\/www.tothenew.com\/blog\/category\/technology\/#listItem","name":"Technology"}}]},{"@type":"Organization","@id":"https:\/\/www.tothenew.com\/blog\/#organization","name":"TO THE NEW Blog","url":"https:\/\/www.tothenew.com\/blog\/"},{"@type":"Person","@id":"https:\/\/www.tothenew.com\/blog\/author\/sachin\/#author","url":"https:\/\/www.tothenew.com\/blog\/author\/sachin\/","name":"Sachin","image":{"@type":"ImageObject","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#authorImage","url":"https:\/\/secure.gravatar.com\/avatar\/d84b0b7a42bd13ad65b8a665e117addb0455be6b706c443a53526dc9d0e5af87?s=96&d=mm&r=g","width":96,"height":96,"caption":"Sachin"}},{"@type":"WebPage","@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#webpage","url":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/","name":"Normalizing Accented Words | TO THE NEW Blog","description":"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like \"saint\" will always","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/www.tothenew.com\/blog\/#website"},"breadcrumb":{"@id":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/#breadcrumblist"},"author":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/sachin\/#author"},"creator":{"@id":"https:\/\/www.tothenew.com\/blog\/author\/sachin\/#author"},"datePublished":"2012-06-14T17:42:24+05:30","dateModified":"2016-12-19T15:16:40+05:30"},{"@type":"WebSite","@id":"https:\/\/www.tothenew.com\/blog\/#website","url":"https:\/\/www.tothenew.com\/blog\/","name":"TO THE NEW Blog","inLanguage":"en-US","publisher":{"@id":"https:\/\/www.tothenew.com\/blog\/#organization"}}]},"og:locale":"en_US","og:site_name":"TO THE NEW BLOG","og:type":"blog","og:title":"Normalizing Accented Words | TO THE NEW Blog","og:description":"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &quot;saint&quot; will always","og:url":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/","og:image":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","og:image:secure_url":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","twitter:card":"summary","twitter:site":"@tothenew","twitter:title":"Normalizing Accented Words | TO THE NEW Blog","twitter:description":"We all often need to work on data aggregated together from different sources, and before we analyse it, we often need to normalize it to a certain standard, A normalization process typically includes removing special characters, converting all text to lower case , We can also have certain rules that words like &quot;saint&quot; will always","twitter:image":"https:\/\/www.tothenew.com\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png"},"aioseo_meta_data":{"post_id":"5754","title":null,"description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"Article","isEnabled":true},"graphs":[]},"schema_type":null,"schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"limit_modified_date":false,"created":"2021-04-29 12:03:31","updated":"2024-02-29 08:18:31","focus_keyword":null,"additional_keywords":null,"truseo_locale":null,"ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.tothenew.com\/blog\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/www.tothenew.com\/blog\/category\/technology\/\" title=\"Technology\">Technology<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tNormalizing Accented Words\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/www.tothenew.com\/blog"},{"label":"Technology","link":"https:\/\/www.tothenew.com\/blog\/category\/technology\/"},{"label":"Normalizing Accented Words","link":"https:\/\/www.tothenew.com\/blog\/normalizing-accented-words\/"}],"_links":{"self":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts\/5754","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/users\/14"}],"replies":[{"embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/comments?post=5754"}],"version-history":[{"count":0,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/posts\/5754\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/media?parent=5754"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/categories?post=5754"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.tothenew.com\/blog\/wp-json\/wp\/v2\/tags?post=5754"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}