As of 2026-09-01 (commit 8dc44a361022), embeddings-benchmark/mteb contains 895,768 total lines: 868,146 code, 5,976 comments, 21,646 blank, across 3,767 files in 11 languages (top: JSON 74.5%). Counted with tokei via OctoCounts.
OctoCounts produced this report by resolving embeddings-benchmark/mteb to commit 8dc44a361022, downloading the repository source archive, and counting every source file with tokei, the open-source line counter written in Rust. The table below breaks the count down by programming language into files, total lines, code lines, comment lines, and blank lines, so the figures can be compared across languages and projects. Results are cached by commit, tokei version, and analysis options, so counting the same revision again reproduces exactly these numbers.
This is a large codebase by counted code lines. Code represents 96.9% of all lines, comments represent 0.7%, and the repository averages 230 code lines per file. JSON accounts for 74.5% of counted code.
| Language | Files | Lines | Code | Comments | Blanks |
|---|---|---|---|---|---|
| JSON | 1932 | 647185 | 647185 | 0 | 0 |
| Python | 1810 | 242554 | 217573 | 3984 | 20997 |
| Jupyter Notebooks | 8 | 3037 | 2259 | 351 | 427 |
| TOML | 1 | 702 | 618 | 41 | 43 |
| Shell | 3 | 307 | 220 | 48 | 39 |
| YAML | 1 | 153 | 140 | 3 | 10 |
| Makefile | 1 | 95 | 71 | 0 | 24 |
| Dockerfile | 1 | 122 | 48 | 47 | 27 |
| HTML | 1 | 33 | 32 | 0 | 1 |
| Markdown | 5 | 324 | 0 | 246 | 78 |
| Plain Text | 4 | 1256 | 0 | 1256 | 0 |
Top language (JSON 74.5%). Generated at 2026-09-01T09:27:42.551616095+00:00.
embeddings-benchmark/mteb has 895,768 total lines, including 868,146 code lines, 5,976 comment lines, and 21,646 blank lines.
OctoCounts resolved the public GitHub repository to commit 8dc44a361022, downloaded the source archive, counted it with tokei, and cached the report by commit, tokei version, and analysis options.
This OctoCounts report was generated from main at commit 8dc44a361022 on 2026-09-01T09:27:42.551616095+00:00.
Other public repositories with an OctoCounts report, ranked by top language and code size similarity: