Senger CodeLab 🚀

How to split a string in Haskell

September 29, 2026

📂 Categories: Programming
How to split a string in Haskell

Navigating string manipulation is a fundamental task in almost any programming language, and Haskell is no exception. While its approach to strings, being lists of characters, offers unique advantages, it also means that common operations like splitting a string might require a slightly different mindset than in imperative languages. If you’ve ever found yourself wondering how to split a string in Haskell, you’re in the right place. This guide will walk you through various methods, from simple whitespace splitting to more advanced techniques using custom delimiters, ensuring you can effectively parse and process textual data in your Haskell applications. We’ll explore core functions, delve into the powerful Data.List.Split module, and even touch upon the differences between lazy Strings and strict Text types, equipping you with the knowledge to handle diverse string splitting challenges with idiomatic Haskell.

Understanding Strings in Haskell: List vs. Text

Before diving into splitting, it’s crucial to understand how Haskell handles strings. By default, a String in Haskell is simply a type alias for a list of characters: type String = [Char]. This fundamental definition means that many list manipulation functions can be directly applied to strings, offering a concise and powerful way to work with them. For instance, concatenating strings is just list concatenation, and taking a substring involves list slicing. However, this list-based approach can become inefficient for very large strings due to the overhead of linked lists, especially when dealing with operations that require traversing the entire string repeatedly.

For performance-critical applications or when processing substantial amounts of text, Haskell programmers often turn to the Text type from the Data.Text module. Unlike String, Text represents strings as arrays of characters (specifically, UTF-16 code units in a packed, efficient format), making many operations, including splitting, significantly faster. Both String and Text have their own sets of functions for manipulation, and it’s common to convert between them using pack and unpack as needed. Throughout this article, we’ll primarily focus on String manipulation for simplicity, but we’ll also highlight equivalent Text functions where relevant, acknowledging the importance of both in the Haskell ecosystem. Understanding these underlying representations is key to choosing the most appropriate method for your specific string processing needs, particularly when parsing complex data structures like CSV files or configuration settings.

Basic String Splitting: The words Function

For the simplest cases, where you need to split a string by any sequence of whitespace characters (spaces, tabs, newlines), Haskell provides a built-in function called words. This function resides in the standard Prelude, meaning you don’t need to import any special modules to use it. It takes a single String as input and returns a list of Strings, where each element is a word from the original string. Consecutive whitespace characters are treated as a single delimiter, and leading/trailing whitespace is ignored, making it ideal for parsing natural language sentences or simple command-line arguments.

Consider the string "Hello world, how are you?". Applying words to this string would yield ["Hello", "world,", "how", "are", "you?"]. Notice that punctuation attached to words remains part of the word, as words only considers whitespace as a separator. This behavior is often desirable for initial tokenization tasks. For more advanced scenarios where punctuation also needs to be separated or specific delimiters are required, you’ll need more powerful tools, but for a quick and easy split by whitespace, words is the go-to function. It’s a testament to Haskell’s design that such a common utility is readily available without additional boilerplate.

Here’s a quick example:

-- Example 1: Basic usage let sentence = " This is a sample sentence with extra spaces. " print (words sentence) -- Output: ["This", "is", "a", "sample", "sentence", "with", "extra", "spaces."] -- Example 2: With newlines and tabs let multiline = "Line1\tLine2\nLine3" print (words multiline) -- Output: ["Line1", "Line2", "Line3"] 

Advanced Splitting with Data.List.Split: The splitOn Function

When the simple words function isn’t sufficient, and you need to split a string by a specific character or even another string, the Data.List.Split module becomes indispensable. This powerful library offers a variety of sophisticated splitting functions, with splitOn being one of the most frequently used. To leverage it, you’ll first need to add split to your project’s dependencies (e.g., in your .cabal file or package.yaml) and then import it: import Data.List.Split.

The splitOn function takes two arguments: the delimiter (which can be a Char or a String) and the string to be split. It returns a list of strings. A key characteristic of splitOn is how it handles delimiters at the beginning or end of the string, or consecutive delimiters. For instance, if the string starts or ends with the delimiter, an empty string will be included at the beginning or end of the result list, respectively. Similarly, consecutive delimiters will result in empty strings between them. This precise behavior allows for consistent parsing, which is critical in scenarios like parsing CSV data where empty fields are meaningful.

For example, to split a string by a comma, you would use splitOn "," "apple,banana,,orange", which yields ["apple","banana","","orange"]. Notice the empty string for the missing field between the two commas. This module also provides other versatile functions like splitEvery for fixed-length chunks, chunksOf, and more, making it a comprehensive toolkit for almost any string splitting requirement. According to Haskell expert and author, Michael Snoyman, “Data.Text and Data.List.Split are two of the most commonly used external libraries for string manipulation in real-world Haskell applications, emphasizing their practical utility.”

Using splitOn with Different Delimiters

The flexibility of splitOn shines when dealing with various delimiter types. Here’s how you can use it effectively:

  1. Splitting by a single character: To split a string by a character like ';', you simply pass the character as the first argument. For example, splitOn ";" "item1;item2;item3" would produce ["item1", "item2", "item3"]. This is common for parsing simple delimited lists.

  2. Splitting by a string: When your delimiter is not just a single character but a sequence of characters, such as "", splitOn handles this gracefully. splitOn "" "part1part2part3" yields ["part1", "part2", "part3"]. This is particularly useful for parsing custom protocols or log file entries.

  3. Handling edge cases: As mentioned, splitOn includes empty strings for delimiters at the start, end, or consecutive Question & Answer :
    Is there a standard way to split a string in Haskell?

    lines and words work great from splitting on a space or newline, but surely there is a standard way to split on a comma?

    I couldn’t find it on Hoogle.

    To be specific, I’m looking for something where split "," "my,comma,separated,list" returns ["my","comma","separated","list"].

    Remember that you can look up the definition of Prelude functions!

    http://www.haskell.org/onlinereport/standard-prelude.html

    Looking there, the definition of words is,

    words :: String -> [String] words s = case dropWhile Char.isSpace s of "" -> [] s' -> w : words s'' where (w, s'') = break Char.isSpace s' 
    

    So, change it for a function that takes a predicate:

    wordsWhen :: (Char -> Bool) -> String -> [String] wordsWhen p s = case dropWhile p s of "" -> [] s' -> w : wordsWhen p s'' where (w, s'') = break p s' 
    

    Then call it with whatever predicate you want!

    main = print $ wordsWhen (==',') "break,this,string,at,commas"