Search Shortcut cmd + k | ctrl + k
markdown

Read, analyze, and write Markdown files with block-level document representation and inline element support

Maintainer(s): teaguesterling

Installing and Loading

INSTALL markdown FROM community;
LOAD markdown;

Example

-- Load the extension
LOAD markdown;

-- Read Markdown files with glob patterns
SELECT content FROM read_markdown('docs/**/*.md');

-- Parse into block-level elements (duck_block shape)
SELECT element_type, content, level
FROM read_markdown_blocks('README.md')
ORDER BY element_order;

-- Extract code blocks from Markdown text
SELECT cb.language, cb.code
FROM (
  SELECT UNNEST(md_extract_code_blocks('```python\nprint("Hello")\n```')) as cb
);

-- Build rich text with inline elements
SELECT duck_blocks_to_md([
  {kind: 'inline', element_type: 'text', content: 'Check out ', level: 1, encoding: 'text', attributes: MAP{}, element_order: 0},
  {kind: 'inline', element_type: 'link', content: 'our docs', level: 1, encoding: 'text', attributes: MAP{'href': 'https://duckdb-markdown.readthedocs.io/'}, element_order: 1}
]);

-- Export query results as Markdown table
COPY (SELECT * FROM my_table) TO 'output.md' (FORMAT MARKDOWN);

-- Round-trip: read blocks, transform, write back
COPY (
  SELECT kind, element_type, content, level, encoding, attributes
  FROM read_markdown_blocks('doc.md')
) TO 'copy.md' (FORMAT MARKDOWN, markdown_mode 'blocks');

About markdown

The Markdown extension adds comprehensive Markdown processing capabilities to DuckDB, enabling structured analysis, transformation, and generation of Markdown documents.

Documentation: https://duckdb-markdown.readthedocs.io/

Key Features:

  • File Reading Functions: Read Markdown files with read_markdown(), read_markdown_sections(), and read_markdown_blocks() supporting glob patterns, metadata extraction, and block-level parsing
  • Block-Level Representation: Parse documents into duck_block format with kind (block/inline), element_type, content, level, encoding, attributes, and element_order columns
  • Inline Element Support: Build rich text content with bold, italic, links, code, math, and more using the unified duck_block structure
  • COPY TO Markdown: Export query results as Markdown tables, documents, or block-level representations with full round-trip support
  • Content Extraction: Extract code blocks, links, images, and tables from Markdown content using structured LIST return types
  • Document Processing: Convert markdown to HTML/text, validate content, extract metadata, and generate document statistics
  • Replacement Scan Support: Query Markdown files directly using FROM '*.md' syntax with full glob pattern support
  • Native MARKDOWN Type: Custom MARKDOWN type with automatic VARCHAR casting for seamless integration
  • Cross-Platform Support: Works on Linux, macOS, WebAssembly, and Windows
  • GitHub Flavored Markdown: Uses cmark-gfm for accurate parsing of modern Markdown features
  • High Performance: Process thousands of documents efficiently with 4,000+ sections/second processing rate

Core Functions:

  • read_markdown() - Read Markdown files with comprehensive parameter support
  • read_markdown_sections() - Parse files into hierarchical sections with filtering options
  • read_markdown_blocks() - Parse files into block-level elements (duck_block shape)
  • duck_block_to_md() - Convert single block/inline element to Markdown
  • duck_blocks_to_md() - Convert list of elements to Markdown document
  • duck_blocks_to_sections() - Convert blocks to hierarchical sections
  • md_extract_code_blocks() - Extract code blocks with language and metadata
  • md_extract_links() - Extract links with text, URL, and title information
  • md_extract_images() - Extract images with alt text and metadata
  • md_extract_tables_json() - Extract tables as structured JSON
  • md_to_html() - Convert markdown content to HTML
  • md_to_text() - Convert markdown to plain text for full-text search
  • md_stats() - Get document statistics (word count, reading time, etc.)
  • md_extract_metadata() - Extract frontmatter metadata as MAP

COPY TO Modes:

  • table (default) - Export any query as a formatted Markdown table
  • document - Reconstruct Markdown from sections with headings and content
  • blocks / duck_block - Round-trip block-level representation with inline element support

Example Use Cases:

  • Documentation analysis across entire repositories
  • Content quality assessment and auditing
  • Large-scale documentation search and indexing
  • Code example extraction and analysis
  • Document transformation pipelines
  • Rich text generation with inline formatting
  • Knowledge base processing and content management

Performance:

Real-world benchmark: Processing 287 Markdown files (2,699 sections, 1,137 code blocks, 1,174 links) in 603ms.

Full test suite with 1948 passing assertions across 54 test files.

Added Functions

| function_name | function_type | description | comment | examples | |——————————-|—————|————————————————————————————|———|———————————————————————————————————————————————————————| | duck_block_to_md | scalar | Convert a single duck_block struct to Markdown text. | NULL | [duck_block_to_md({'kind': 'container', 'element_type': 'paragraph', 'content': 'hello', 'level': 0, 'encoding': 'text', 'attributes': map(), 'element_order': 0})] | | duck_blocks_to_md | scalar | Convert a list of duck_blocks to Markdown text. | NULL | [duck_blocks_to_md([])] | | duck_blocks_to_sections | scalar | Convert a list of duck_blocks into structured sections. | NULL | [duck_blocks_to_sections([])] | | md_extract_code_blocks | scalar | Extract fenced and indented code blocks from Markdown content. | NULL | [md_extract_code_blocks('sql SELECT 1; ')] | | md_extract_frontmatter | scalar | Extract the raw YAML frontmatter text block from Markdown. | NULL | [md_extract_frontmatter('— title: Test —

Body')] |

| md_extract_images | scalar | Extract images from Markdown content. | NULL | [md_extract_images('Logo')] | | md_extract_links | scalar | Extract hyperlinks from Markdown content. | NULL | [md_extract_links('DuckDB')] | | md_extract_metadata | scalar | Extract YAML frontmatter from Markdown text as a MAP. | NULL | [md_extract_metadata('— title: Test —

Body')] |

| md_extract_section | scalar | NULL | NULL | | | md_extract_sections | scalar | NULL | NULL | | | md_extract_table_rows | scalar | Extract table cell rows from Markdown content. | NULL | [md_extract_table_rows('| a | b | |—|—| | 1 | 2 |')] | | md_extract_tables_json | scalar | Extract tables from Markdown content as structured objects. | NULL | [md_extract_tables_json('| a | b | |—|—| | 1 | 2 |')] | | md_extract_tags | scalar | Extract hashtag tags from Markdown content. | NULL | [md_extract_tags('Tag #important text')] | | md_extract_wikilinks | scalar | Extract wiki-style links and embeds from Markdown content. | NULL | [md_extract_wikilinks('[[Page Name]]')] | | md_section_breadcrumb | scalar | Generate a breadcrumb string combining file path and section identifier. | NULL | [md_section_breadcrumb('doc.md', 'intro')] | | md_stats | scalar | NULL | NULL | | | md_to_html | scalar | Convert Markdown text to HTML. | NULL | [md_to_html('# Title')] | | md_to_text | scalar | Convert Markdown text to plain text. | NULL | [md_to_text('# Title')] | | md_valid | scalar | Validate markdown content. | NULL | [md_valid('# Hello')] | | parse_markdown_to_duck_blocks | scalar | Parse Markdown text into a list of canonical duck_block structs. | NULL | [parse_markdown_to_duck_blocks('# Hello')] | | read_markdown | table | Read Markdown documents from files into table format with frontmatter and content. | NULL | [SELECT * FROM read_markdown('README.md')] | | read_markdown_blocks | table | Read Markdown files parsed into atomic block elements. | NULL | [SELECT * FROM read_markdown_blocks('README.md')] | | read_markdown_sections | table | Read Markdown files split by headings into structured sections. | NULL | [SELECT * FROM read_markdown_sections('README.md')] | | value_to_md | scalar | Convert any value to a Markdown formatted string. | NULL | [value_to_md(42)] |

Overloaded Functions

This extension does not add any function overloads.

Added Types

type_name type_size logical_type type_category internal
markdown 16 VARCHAR STRING true
md 16 VARCHAR STRING true

Added Settings

This extension does not add any settings.