ClickHouse HOLs
GitHub

ClickHouse 26.2 New Features Lab


A hands-on laboratory for learning and testing ClickHouse 26.2 new features. This directory focuses on verified and working features newly added in ClickHouse 26.2 (released 2026-02-26).

📋 Overview

ClickHouse 26.2 expands the SQL surface area with number-theory utilities, a new 128-bit hash function, and an introspectable text-index tokenizer catalog. The release also makes the text index GA, promotes the QBit vector type to GA, and adds distributed vector search.

🎯 Key Features

  1. primes() Table Function + system.primes — prime numbers as a first-class SQL source
  2. xxh3_128 Hash Function — 128-bit non-cryptographic hash for deduplication, sharding, and joins
  3. system.tokenizers Catalog — discover every tokenizer available to text indexes from SQL

🚀 Quick Start

Prerequisites

Setup and Run

# 1. Install and start ClickHouse 26.2
cd local/releases/26.2
./00-setup.sh

# 2. Run feature tests
./01-primes-function.sh
./02-xxh3-128-hash.sh
./03-system-tokenizers.sh

Manual Execution (SQL only)

cd ../../oss-docker
./client.sh 8123

cd ../local/releases/26.2
source 01-primes-function.sql

📚 Feature Details

1. primes() Table Function & system.primes (01-primes-function)

New Feature: A new table function primes(n) returns the first n prime numbers, and system.primes exposes an unbounded stream of primes — a single column named prime of type UInt64. (#92776)

Test Content: - Reading from system.primes with LIMIT and WHERE - Generating bounded prime sequences via primes(n) - Twin-prime pair detection with self-JOIN on row number - Prime-gap distribution analysis - Composite-number identification via NOT IN (primes) - Building a materialized prime cache table - Nearest-prime lookup for hash-table sizing

Key Learning Points: - Both primes() and system.primes use a single column named prime (not number) - system.primes is unbounded — always pair with LIMIT or a WHERE prime < X predicate - primes(n) is bounded and is the safer choice when you know the upper bound - Combine with rowNumberInAllBlocks() to access ordinal position

Use Cases: - Hash-table bucket sizing (use the next prime ≥ desired capacity) - Modular arithmetic and number-theory experiments - Generating test/synthetic data with deterministic distinctness - Teaching and demonstrations


2. xxh3_128 Hash Function (02-xxh3-128-hash)

New Feature: 128-bit XXH3 hash returning UInt128, doubling the hash space of the existing xxHash64 at comparable throughput. (#96055)

Test Content: - Basic hashing and UInt128 result type inspection - Hex / decimal representations side-by-side with xxHash64 - Content-fingerprint deduplication via MATERIALIZED column - Hash-based shard mapping - Collision test on 10,000 random strings - JOIN on 128-bit fingerprints

Key Learning Points: - Result type is UInt128 (32 hex chars vs xxHash64's 16) - Same input → same hash (deterministic, no salt) - Use as a MATERIALIZED column for cheap, indexable content fingerprints - Practical collision probability is negligible at typical data sizes - The function name is xxh3_128, while the existing 64-bit equivalent is xxHash64 (note the camel-case asymmetry)

Use Cases: - Deduplication where 64-bit collision risk is unacceptable - Sharding keys for very large key spaces - Content-addressable joins across tables - Idempotency keys / cache lookups


3. system.tokenizers Catalog (03-system-tokenizers)

New Feature: system.tokenizers system table lists every tokenizer registered for use with text indexes and the tokens() function. (#96753)

Test Content: - Listing all tokenizers from system.tokenizers - Calling tokens(s, tokenizer_name [, params]) with different tokenizers (splitByNonAlpha, splitByString, ngrams) - Creating a text index with an explicit tokenizer - Full-text search via hasToken driven by the index - Token count per document - EXPLAIN indexes = 1 to confirm the skip index is used - Side-by-side comparison of tokenizer output

Key Learning Points: - Available tokenizers include splitByNonAlpha, splitByString, ngrams, ngrambf_v1, tokenbf_v1, sparseGrams, sparse_grams, array - text indexes are now GA in 26.2 (no compatibility flag required) - tokens(string, tokenizer_name) lets you preview what the index will store before building it - Tokenizer choice directly determines what queries the index can accelerate

Use Cases: - Choosing the right tokenizer before building a large text index - Validating tokenization behavior against multilingual or structured text - Diagnosing why a text-index query is or isn't being skipped


🔧 Management

ClickHouse Connection Info

Useful Commands

cd ../../oss-docker
./status.sh
./client.sh 8123
docker logs clickhouse-26-2
./stop.sh
./stop.sh --cleanup

📂 File Structure

26.2/
├── README.md                     # This document
├── 00-setup.sh                   # ClickHouse 26.2 installation script
├── 01-primes-function.sh         # primes()/system.primes test runner
├── 01-primes-function.sql        # primes()/system.primes SQL
├── 02-xxh3-128-hash.sh           # xxh3_128 test runner
├── 02-xxh3-128-hash.sql          # xxh3_128 SQL
├── 03-system-tokenizers.sh       # system.tokenizers test runner
└── 03-system-tokenizers.sql      # system.tokenizers SQL

🆕 What's New in 26.2

🔍 Additional Resources

📝 Notes

📄 License

MIT — free to learn from and modify.


Happy Learning! 🚀

For questions or issues, see the main clickhouse-hols README.


ClickHouse 26.2 신기능 테스트 및 학습 환경입니다. 이 디렉토리는 2026년 2월 26일 출시된 ClickHouse 26.2에서 검증된 작동 기능에 집중합니다.

📋 개요

ClickHouse 26.2는 수론 유틸리티, 새로운 128비트 해시 함수, 그리고 검색 가능한 텍스트 인덱스 토크나이저 카탈로그로 SQL 표면을 확장합니다. 또한 텍스트 인덱스가 GA로 승격되고, QBit 벡터 타입이 GA가 되었으며, 분산 벡터 검색이 추가되었습니다.

🎯 주요 기능

  1. primes() 테이블 함수 + system.primes — 소수를 SQL 소스로 직접 활용
  2. xxh3_128 해시 함수 — 중복 제거·샤딩·조인을 위한 128비트 비암호 해시
  3. system.tokenizers 카탈로그 — 텍스트 인덱스용 토크나이저 목록을 SQL로 조회

🚀 빠른 시작

사전 요구사항

설정 및 실행

# 1. ClickHouse 26.2 설치 및 시작
cd local/releases/26.2
./00-setup.sh

# 2. 각 기능별 테스트 실행
./01-primes-function.sh
./02-xxh3-128-hash.sh
./03-system-tokenizers.sh

📚 기능 상세

1. primes() 테이블 함수 & system.primes

새 테이블 함수 primes(n)은 첫 n개의 소수를 반환하고, system.primes는 무한 소수 스트림을 제공합니다. 컬럼명은 prime(UInt64) 하나입니다. (#92776)

테스트 내용: - LIMIT/WHERE로 system.primes 조회 - primes(n)로 유한 소수 시퀀스 생성 - 행 번호 자기 조인으로 쌍둥이 소수 탐색 - 소수 간격(prime gap) 분포 분석 - NOT IN (primes)로 합성수 식별 - 소수 캐시 테이블 구축 - 해시 테이블 크기용 최근접 소수 검색

2. xxh3_128 해시 함수

UInt128을 반환하는 128비트 XXH3 해시 — 기존 xxHash64의 해시 공간을 두 배로 확장하면서 비슷한 처리량을 유지합니다. (#96055)

테스트 내용: - 기본 해싱 및 UInt128 결과 타입 확인 - xxHash64와의 16진수/십진수 표현 비교 - MATERIALIZED 컬럼을 통한 콘텐츠 지문 기반 중복 제거 - 해시 기반 샤드 매핑 - 10,000개 랜덤 문자열 충돌 테스트 - 128비트 지문 기반 JOIN

3. system.tokenizers 카탈로그

system.tokenizers는 text 인덱스와 tokens() 함수에서 사용 가능한 모든 토크나이저를 보여줍니다. (#96753)

테스트 내용: - system.tokenizers에서 토크나이저 목록 조회 - 다양한 토크나이저로 tokens(s, name [, params]) 호출 - 명시적 토크나이저로 text 인덱스 생성 - 인덱스를 활용한 hasToken 전체 텍스트 검색 - 문서별 토큰 개수 계산 - EXPLAIN indexes = 1로 스킵 인덱스 사용 확인

🆕 26.2의 새로운 기능

🔍 추가 자료

📝 참고사항

📄 라이선스

MIT — 자유롭게 학습하고 수정하세요.


Happy Learning! 🚀

질문이나 이슈는 메인 clickhouse-hols README를 참조하세요.

Open this lab on GitHub →GitHub에서 이 실습 열기 →