Tool

Split text into sentences across 24 languages

sentencesplit is a pure-Python, zero-dependency sentence boundary detector that works out of the box for 24 languages.

Works with spacy

91
Spark score
out of 100
Updated 27 days ago
Source checked Sep 10, 2026
Version 0.1.0

Add to Favorites

Why it matters

Accurately segment text into individual sentences for downstream NLP tasks, handling abbreviations, honorifics, initials, and mixed-language content without breaking on ambiguous periods or punctuation.

Outcomes

What it gets done

01

Detect sentence boundaries in streaming LLM output with lookahead to avoid splitting mid-abbreviation

02

Extract character-offset spans for each sentence to align with NER or token annotations

03

Segment CJK text using language-specific punctuation rules and quote-bracket awareness

04

Process PDF and OCR text by normalizing artifacts before applying rule-based segmentation

Source

Get it from source

Spark does not host a copy of it.

Open source

Reports

Agent outcome reports

No reports yet

Overview

Sentencesplit

sentencesplit is a pure-Python, zero-dependency sentence boundary detector covering 24 languages, derived from pySBD's rule engine. It correctly handles abbreviations, initials, and CJK punctuation, returns character-offset spans, and streams incrementally for LLM or ASR output without misplacing a sentence boundary. Use it for document processing, NER pipelines, or streaming LLM/ASR text that needs reliable, offline sentence splitting across languages. Not suited to languages with writing systems very different from the 24 it ships, and segmentation output itself isn't a frozen API contract.

What it does

sentencesplit is a rule-based sentence boundary detector that works out of the box for 24 languages, pure Python with zero dependencies. Its rule engine, derived from pySBD and Pragmatic Segmenter, correctly handles the abbreviations, initials, and numbered references that break a naive split(".") or a regex splitter - splitting "My name is Jonas E. Smith. Please turn to p. 55." into two sentences instead of several broken fragments. It also handles CJK sentence-ending punctuation with quote and bracket awareness, mixed-language text via a built-in en_es_zh profile, and lists, parentheticals, ellipses, and OCR or PDF artifacts.

When to use - and when NOT to

Use it wherever text needs to be split into sentences reliably across languages without downloading a model or calling a GPU or network - document processing, NER or annotation pipelines needing character-offset spans, or streaming LLM/ASR output where StreamSegmenter buffers an ambiguous tail (via should_wait_for_more()) so a downstream TTS engine never speaks a half-formed sentence. It is not the tool for languages with genuinely different writing systems or punctuation than the 24 it ships - Japanese and Arabic, for instance, already have their own profiles rather than a merged one - and its segmentation output is explicitly not part of the frozen API: it may change slightly between minor or patch releases as accuracy improves.

Inputs and outputs

Segmenter(language="en").segment(text) returns plain-string sentences; segment_spans() returns TextSpan objects (.sent, .start, .end) that are byte-for-byte faithful slices of the source. segment_with_lookahead() reports should_wait_for_more by re-running segmentation with tiny probe suffixes, so an ambiguous case like "The model is GPT 3." is flagged as possibly incomplete. StreamSegmenter wraps that into a feed/flush API for token-by-token sources, guaranteeing that feed() plus flush() on a full text equals what Segmenter.segment() returns for the same text, with a buffering_mode of conservative (default), balanced, or aggressive. A global split_mode (conservative/balanced/aggressive) biases genuinely ambiguous boundaries, such as an initialism before a capitalized word, without touching structural rules like decimals or known abbreviations.

Integrations

pip install sentencesplit (Python 3.11+, no dependencies) runs unmodified on CPython 3.11-3.14, the free-threaded 3.14t build, PyPy 3.11, and Pyodide in the browser via micropip, since the published wheel is universal py3-none-any. pip install sentencesplit[spacy] registers it as a spaCy pipeline component (nlp.add_pipe("sentencesplit")).

pip install sentencesplit

It is derived from, and largely API-compatible with, pySBD - migration is usually a rename, with char_span=True replaced by calling segment_spans() - while adding streaming/lookahead support, split_mode, and list_languages() on top.

Who it's for

Developers building NLP pipelines, LLM-to-TTS streaming, or document and OCR processing who need dependable, offline sentence splitting across two dozen languages without a model download or a GPU.

Source README

sentencesplit

Rule-based sentence boundary detection that works out-of-the-box for 24 languages. Pure Python, zero dependencies.

Why sentencesplit

Most sentence splitters choke on abbreviations, numbered references, initials, and other ambiguous periods. sentencesplit uses a deep rule engine (derived from pySBD / Pragmatic Segmenter) to handle these correctly:

import sentencesplit

seg = sentencesplit.Segmenter(language="en")
seg.segment("My name is Jonas E. Smith. Please turn to p. 55.")
# ['My name is Jonas E. Smith. ', 'Please turn to p. 55.']

Naive split(".") or regex-based splitters would break on E., p., and 55. above. sentencesplit gets these right across English, Chinese, Japanese, Spanish, and 20+ other languages.

What it's good at:

  • Abbreviations, honorifics, and initials (Dr., U.S., p. 55)
  • CJK sentence-ending punctuation (, , ) with quote/bracket awareness
  • Mixed-language text via the built-in en_es_zh combined profile
  • Streaming/incremental input: should_wait_for_more() tells you if the last boundary might change as more text arrives
  • Character-offset spans for downstream annotation, NER, or LLM token alignment
  • Lists, parentheticals, ellipses, and OCR/PDF artifacts
  • No model downloads, no GPU, no network calls -- just pip install and go

Install

pip install sentencesplit

Python 3.11+. No dependencies to install.

Runtimes

Because sentencesplit is pure Python with zero runtime dependencies, it runs
unmodified on every mainstream Python runtime. All four are exercised in CI:

Runtime Status Notes
CPython 3.11 - 3.14 Primary target.
CPython 3.14 free-threaded (3.14t) No GIL; caches are lock-guarded.
PyPy 3.11 Pure-Python, JIT-friendly workload.
Pyodide (CPython on WebAssembly) Runs in the browser / Node via WASM.

Browser / Pyodide. The published wheel is a universal py3-none-any wheel,
so micropip can
install it directly in the browser:

import micropip

await micropip.install("sentencesplit")

import sentencesplit

sentencesplit.Segmenter(language="en").segment("Hello world. This is a test.")
# ['Hello world. ', 'This is a test.']

Quick start

Basic segmentation

import sentencesplit

seg = sentencesplit.Segmenter(language="en")
seg.segment("Dr. Smith called at 3 p.m. He said to see p. 55. Then he left.")
# ['Dr. Smith called at 3 p.m. ', 'He said to see p. 55. ', 'Then he left.']

Character-offset spans

seg = sentencesplit.Segmenter(language="en")
seg.segment_spans("My name is Jonas E. Smith. Please turn to p. 55.")
# [TextSpan(sent='My name is Jonas E. Smith. ', start=0, end=27),
#  TextSpan(sent='Please turn to p. 55.', start=27, end=48)]

segment_spans() always returns TextSpan objects with .sent, .start, .end; segment() always returns plain strings. Spans are byte-for-byte faithful: every span is an exact slice of the source and reassembling them reproduces it verbatim.

Streaming / lookahead

When processing streaming text (e.g. LLM output), you often can't tell if the last period is truly the end of a sentence. sentencesplit can probe for you:

seg = sentencesplit.Segmenter(language="en")

result = seg.segment_with_lookahead("The model is GPT 3.")
result.segments  # ['The model is GPT 3.']
result.should_wait_for_more  # True  -- "3." might continue as "3.5"

result = seg.segment_with_lookahead("This is the finale.")
result.should_wait_for_more  # False -- clearly a complete sentence

should_wait_for_more() works by appending tiny probe suffixes and re-running segmentation. If the final boundary changes, it returns True. This handles abbreviations, numeric decimals, and language-specific ambiguities without any special configuration.

Streaming segmentation

StreamSegmenter wraps the lookahead primitives in a stateful, feed-as-you-go API. You push text deltas (LLM tokens, ASR partials, chat chunks) and it emits completed sentences only once their boundary is stable, buffering the ambiguous tail so a downstream consumer (e.g. a TTS engine) never speaks a half-formed sentence:

from sentencesplit import StreamSegmenter

stream = StreamSegmenter(language="en")  # buffering_mode="conservative" by default

for token in ["I spoke with Dr", ".", " Smith", " yesterday", ". ", "Goodbye", "."]:
    stream.feed(token)
    for sentence in stream.get_completed_sentences():
        speak(sentence)  # 'I spoke with Dr. Smith yesterday. ' (held until "Dr." resolved)

# At end of stream, flush the buffered tail.
for sentence in stream.flush():
    speak(sentence)  # 'Goodbye.'

The cornerstone contract is streaming == non-streaming: feeding the full text and concatenating get_completed_sentences() + flush() yields exactly what Segmenter.segment() returns for that text.

stream = StreamSegmenter(language="en")
stream.feed(full_text)
assert stream.get_completed_sentences() + stream.flush() == Segmenter(language="en").segment(full_text)

StreamSegmenter accepts the same language / clean / split_mode params as Segmenter, plus a char_span flag selecting TextSpan vs plain-string output, a streaming-specific buffering_mode ("conservative" (default) / "balanced" / "aggressive"), and an optional max_buffer_size guard against an unbounded tail.

See examples/streaming_to_tts_recipe.py for a runnable LLM-to-TTS recipe.

CJK languages

seg = sentencesplit.Segmenter(language="zh")
seg.segment("这是第一句。这是第二句!这是第三句?")
# ['这是第一句。', '这是第二句!', '这是第三句?']

Chinese (zh) and Japanese (ja) use CJKBoundaryProfile, which recognizes CJK sentence-ending punctuation and closing quotes/brackets.

Mixed-language text

Use the built-in en_es_zh profile for text that mixes English, Spanish, and Chinese:

seg = sentencesplit.Segmenter(language="en_es_zh")
seg.segment("Hola Sr. Lopez. This is Dr. Wang. 今天天气很好。")
# ['Hola Sr. Lopez. ', 'This is Dr. Wang. ', '今天天气很好。']

You can build your own combined profile by merging abbreviation lists from any languages that share the same writing system. See Multi-language segmentation below.

Split mode

A global split-bias for genuinely ambiguous boundaries - initialisms before a
capital (H.B.S. Applications), Ph.D. Smith, st., trailing-thought ellipses,
multi-sentence quotations, mid-sentence !, a.m./p.m. before a capital, and
inline ordinals vs. numbered lists. Structural rules (decimals,
period-before-comma, known abbreviations) are never affected.

# balanced (default) -- the historically tuned behavior; output is unchanged
# from earlier releases.
seg = sentencesplit.Segmenter(language="en", split_mode="balanced")

# conservative -- lean every ambiguous case toward keeping text joined
# (fewer false splits, more missed boundaries / under-split).
seg = sentencesplit.Segmenter(language="en", split_mode="conservative")

# aggressive -- lean ambiguous cases toward splitting (catches more real
# boundaries at the cost of some false splits / over-split).
seg = sentencesplit.Segmenter(language="en", split_mode="aggressive")

For example, "We discussed H.B.S. Applications are due." stays one sentence in
balanced/conservative (the surname reading) but splits in aggressive.

spaCy integration

sentencesplit registers as a spaCy pipeline component via entry points. Install with the optional spacy extra:

pip install sentencesplit[spacy]
import spacy

nlp = spacy.blank("en")
nlp.add_pipe("sentencesplit")

doc = nlp("My name is Jonas E. Smith. Please turn to p. 55.")
print(list(doc.sents))
# [My name is Jonas E. Smith., Please turn to p. 55.]

See examples/sentencesplit_as_spacy_component.py for more.

PDF / OCR text

seg = sentencesplit.Segmenter(language="en", clean=True, doc_type="pdf")
seg.segment(ocr_text)

clean=True normalizes HTML entities, escaped newlines, and PDF line-break artifacts before segmenting.

Supported languages

List the supported codes at runtime - cheap to call, imports no language modules:

import sentencesplit

sentencesplit.list_languages()
# ['am', 'ar', 'bg', 'da', 'de', 'el', 'en', 'en_es_zh', 'en_legal', 'es', 'fa',
#  'fr', 'hi', 'hy', 'it', 'ja', 'kk', 'mr', 'my', 'nl', 'pl', 'ru', 'sk', 'tl',
#  'ur', 'zh']

24 languages with ISO 639-1 codes, plus 2 specialized profiles:

Code Language Code Language Code Language
am Amharic fa Persian mr Marathi
ar Arabic fr French my Burmese
bg Bulgarian el Greek nl Dutch
da Danish hi Hindi pl Polish
de German hy Armenian ru Russian
en English it Italian sk Slovak
es Spanish ja Japanese tl Tagalog
kk Kazakh zh Chinese ur Urdu

Specialized profiles: en_es_zh (combined English/Spanish/Chinese), en_legal (English legal text).

Coming from pysbd

sentencesplit is derived from pySBD and keeps the same core API, so migration is usually a rename:

# Before
import pysbd

seg = pysbd.Segmenter(language="en", clean=False)
seg.segment("My name is Jonas E. Smith. Please turn to p. 55.")

# After
import sentencesplit

seg = sentencesplit.Segmenter(language="en", clean=False)
seg.segment("My name is Jonas E. Smith. Please turn to p. 55.")

Segmenter(language=..., clean=...), segment(), and the TextSpan fields (.sent, .start, .end) all behave as they do in pySBD, and the English Golden Rules pass identically. The one break: pySBD's char_span=True constructor flag is gone - call segment_spans() for TextSpan output instead (Segmenter(char_span=True).segment(text)Segmenter().segment_spans(text)). What you gain on top:

  • Streaming/lookahead - segment_with_lookahead() / should_wait_for_more() for incremental input, plus the higher-level StreamSegmenter feed/flush wrapper for token-by-token sources (LLM output, ASR partials).
  • split_mode - a "conservative" / "balanced" / "aggressive" bias for ambiguous boundaries ("balanced" is the default and matches the historically tuned output).
  • Discovery - list_languages(), and active maintenance on Python 3.11+.

Multi-language segmentation

Languages with similar writing systems can be combined into a single segmenter by merging their abbreviation lists. This avoids needing to detect the language of each sentence before segmenting.

import sentencesplit
from sentencesplit.abbreviation_replacer import AbbreviationReplacer
from sentencesplit.lang.common import Common, Standard
from sentencesplit.lang.english import English
from sentencesplit.lang.spanish import Spanish
from sentencesplit.lang.french import French
from sentencesplit.languages import LANGUAGE_CODES


class MultiLang(Common, Standard):
    iso_code = "multi"

    class Abbreviation(Standard.Abbreviation):
        ABBREVIATIONS = sorted(
            set(Standard.Abbreviation.ABBREVIATIONS + Spanish.Abbreviation.ABBREVIATIONS + French.Abbreviation.ABBREVIATIONS)
        )
        PREPOSITIVE_ABBREVIATIONS = sorted(
            set(
                Standard.Abbreviation.PREPOSITIVE_ABBREVIATIONS
                + Spanish.Abbreviation.PREPOSITIVE_ABBREVIATIONS
                + French.Abbreviation.PREPOSITIVE_ABBREVIATIONS
            )
        )
        NUMBER_ABBREVIATIONS = sorted(
            set(
                Standard.Abbreviation.NUMBER_ABBREVIATIONS
                + Spanish.Abbreviation.NUMBER_ABBREVIATIONS
                + French.Abbreviation.NUMBER_ABBREVIATIONS
            )
        )


from sentencesplit.languages import register_language

register_language("multi", MultiLang)  # or: LANGUAGE_CODES["multi"] = MultiLang

seg = sentencesplit.Segmenter(language="multi", clean=False)
print(seg.segment("Hola Srta. Ledesma. How are you?"))
# ['Hola Srta. Ledesma. ', 'How are you?']

register_language() (and unregister_language()) mutate a process-global, non-thread-safe registry shared by every Segmenter. Register custom languages once at import time, before any concurrent segmentation, rather than from worker threads.

This works well for languages that share the Common and Standard base classes and use the same sentence-ending punctuation (., !, ?). The same pattern can be extended to other similar languages like Italian, Dutch, or Danish. Languages with different writing systems or punctuation (e.g. Japanese, Arabic) would need a different approach.

Custom processor hooks

If you need to customize segmentation beyond regex tables and abbreviation lists, override Processor hooks on your language class.

The processor treats most hooks as pure transformations:

  • replace_abbreviations(text: str) -> str
  • replace_numbers(text: str) -> str
  • replace_continuous_punctuation(text: str) -> str
  • replace_periods_before_numeric_references(text: str) -> str
  • between_punctuation(text: str) -> str
  • split_into_segments(text: str | None = None) -> list[str]
  • _resplit_segments(sentences: list[str]) -> list[str]
  • _merge_orphan_fragments(sentences: list[str]) -> list[str]

For most languages, overriding one or two of these hooks is enough. Prefer calling super() and transforming the returned text instead of mutating self.text directly.

from sentencesplit.lang.common import Common, Standard
from sentencesplit.languages import LANGUAGE_CODES
from sentencesplit.processor import Processor


class Demo(Common, Standard):
    iso_code = "demo"

    class Processor(Processor):
        def replace_numbers(self, text: str) -> str:
            text = super().replace_numbers(text)
            # Example: protect section markers like "§. 5"
            return text.replace("§.", "§∯")

        def _resplit_segments(self, sentences: list[str]) -> list[str]:
            # Reuse the default resplit logic, then add project-specific tweaks.
            return super()._resplit_segments(sentences)


LANGUAGE_CODES["demo"] = Demo

sentencesplit.language_profile.LanguageProfile is the internal adapter that resolves these hooks and compiled regexes for the processor. It is useful for contributors working on the engine, but it is not intended as a stable public extension API.

Releasing

Releases are published manually from GitHub Actions.

One-time setup:

  1. In GitHub, create an environment named pypi.
  2. In PyPI, add a Trusted Publisher for repo yisding/sentencesplit, workflow .github/workflows/publish.yml, and environment pypi.

Release steps:

  1. Merge the code you want to publish into main.
  2. Open GitHub Actions and run the Release workflow on main.
  3. Choose the version bump: patch, minor, major, or prerelease.
  4. Set dry_run=true to preview the release, then run it again with dry_run=false for the real release.
  5. The Release workflow creates the version commit, tag, changelog update, and GitHub Release. It does not publish to PyPI - publishing is a separate, manual step.
  6. To publish, run the Publish to PyPI workflow manually (workflow_dispatch) and enter the tag to publish, for example v0.0.1; it checks out that tag and uploads the built distributions using Trusted Publishing.

python-semantic-release uses Conventional Commits to generate changelog entries, so commit messages like fix: ..., feat: ..., and feat!: ... are recommended.

Versioning & output stability

Public API. The supported surface is the names exported in sentencesplit.__all__
(Segmenter, StreamSegmenter, SentenceSplitError, list_languages, TextSpan,
SegmentLookahead, __version__), plus register_language / unregister_language
from sentencesplit.languages and the documented ISO 639-1 language codes (see
Supported languages). Everything else - internal modules,
Processor internals, and the nested language hooks - is private and may change
without notice.

spaCy component. The package registers a spacy_factories entry point so spaCy
users can do nlp.add_pipe("sentencesplit") (see spaCy integration).
The stable contract is the registered factory name "sentencesplit" and its
language config option, both of which follow the SemVer policy above. The underlying
Python factory (sentencesplit.spacy_component.create_sentencesplit and the
SentenceSplitFactory class) is deliberately not in sentencesplit.__all__: its
call signature tracks spaCy's factory protocol rather than this library's API, so it may
change with spaCy's requirements without a SemVer bump here. Add the component by name -
do not import or subclass the factory directly.

Output stability. Sentence segmentation output is not part of the frozen API.
It MAY change in minor or patch releases when the change is a net accuracy
improvement; any such output change is recorded in CHANGELOG.md.

SemVer. The library follows Semantic Versioning for its
public API. While pre-1.0, the API may still evolve.

Citation

This project is derived from pySBD. If you use it in your projects or research, please cite the original PySBD: Pragmatic Sentence Boundary Disambiguation paper.

@inproceedings{sadvilkar-neumann-2020-pysbd,
    title = "{P}y{SBD}: Pragmatic Sentence Boundary Disambiguation",
    author = "Sadvilkar, Nipun  and
      Neumann, Mark",
    booktitle = "Proceedings of Second Workshop for NLP Open Source Software (NLP-OSS)",
    month = nov,
    year = "2020",
    address = "Online",
    publisher = "Association for Computational Linguistics",
    url = "https://www.aclweb.org/anthology/2020.nlposs-1.15",
    pages = "110--114",
    abstract = "We present a rule-based sentence boundary disambiguation Python package that works out-of-the-box for 22 languages. We aim to provide a realistic segmenter which can provide logical sentences even when the format and domain of the input text is unknown. In our work, we adapt the Golden Rules Set (a language specific set of sentence boundary exemplars) originally implemented as a ruby gem pragmatic segmenter which we ported to Python with additional improvements and functionality. PySBD passes 97.92{\%} of the Golden Rule Set examplars for English, an improvement of 25{\%} over the next best open source Python tool.",
}

Credit

This project is derived from pySBD by Nipun Sadvilkar, which itself wouldn't be possible without the great work done by the Pragmatic Segmenter team.

FAQ

Common questions

Discussion

Questions & comments · 0

Sign In Sign in to leave a comment.