Run Unix commands with structured XML/JSON for AI agents
Unix coreutils reimplemented with structured XML/JSON output for AI agents - labeled fields, absolute paths, auto-detected language and MIME type.
2.2.0Add to Favorites
Why it matters
Enable AI agents to execute file system operations, text processing, and system queries with machine-readable structured output instead of parsing human-formatted plaintext, eliminating token waste on column guessing and format ambiguity.
Outcomes
What it gets done
List directories with automatic language detection, MIME types, and absolute paths in XML/JSON
Search files with grep, find, and diff returning structured results with line numbers and context
Process text with sort, uniq, cut, sed, and awk outputting parseable data structures
Expose 33 Unix tools as MCP server functions callable directly from Claude Desktop and Claude Code
Source
Get it from source
Spark does not host a copy of it.
Open sourceReports
Agent outcome reports
No reports yet
Overview
Aict
aict reimplements 34 Unix coreutils (ls, grep, cat, diff, find, sed, awk, and more) with structured XML or JSON output instead of plaintext, for AI agents to consume without parsing ambiguity - every field labeled, paths absolute, language and MIME type auto-detected. It ships a built-in MCP server (aict mcp) exposing every tool as a callable function. Use it when an agent chains multiple filesystem lookups and needs unambiguous, structured output; the project's own benchmark shows it costs 1.1-7.8x more tokens per task than plain GNU output but reduces round-trips and parsing errors - use --plain when only raw content is needed.
What it does
aict reimplements 34 Unix coreutils (ls, grep, cat, diff, find, sed, awk, and more) with structured XML or JSON output instead of column-aligned plaintext, specifically for AI agents to consume without parsing ambiguity. Every field is labeled, paths are always absolute, timestamps are Unix integers, and language/MIME type are auto-detected - so an agent doesn't have to guess which column is the file size or whether an entry is a directory.
When to use - and when NOT to
Use it when an agent needs to chain multiple lookups off one command's output - the project's own honest benchmark shows aict output costs 1.1-7.8x more tokens per task than terse GNU output (measured with tiktoken's o200k_base), but that trade buys fewer round-trips: finding .go files with size and modification time takes 4 chained GNU calls (47 tokens) versus 1 aict call (367 tokens, 7.8x), while listing a directory with size/type/language goes from 2 GNU calls to 1 aict call (246 to 820 tokens, 3.3x). In a live agent evaluation (opencode, 3 runs per toolchain on the same task), the aict-equipped agent produced about 46% fewer output tokens overall (median 265 vs 487) and was correct all 3 times, while the GNU-equipped agent shipped a flawed report in the run where it trusted the classic file(1) command's language detection - which, in the same benchmark, misidentified a Go source file as "C source" where aict correctly labeled it go. It also trades raw speed for that semantic richness: --xml mode runs roughly 2-100x slower than the bare GNU tool depending on the operation (grep on 100k lines: 1.3ms GNU vs 130ms aict, 96x, since it's Go's regexp engine against GNU's SIMD-optimized grep). Use --plain when you only need raw content and don't want the enrichment overhead or the extra tokens.
Inputs and outputs
Input is the same arguments a human would pass to the underlying Unix tool - a path for ls/cat/stat, a pattern for grep/find, two files for diff. Output is XML by default (denser than JSON in a context window - <file size="1024" lang="go"/> versus the equivalent JSON object), or --json/--plain on request; a compact mode (the default) uses short attribute names (p/a for path/absolute, t/ma for timestamp/modified-ago) to save tokens further, with --no-compact for verbose names and --dict to see the mapping. Empty results always return valid XML with zero counts rather than an error, and unsupported operations on a given platform return a structured <error> element instead of crashing.
Integrations
{
"mcpServers": {
"aict": {
"command": "aict",
"args": ["mcp"]
}
}
}
aict mcp is a subcommand of the main binary that exposes every tool as a callable MCP function over stdio - no shell wrapping needed - and drops straight into Claude Desktop's or Claude Code's MCP config. It installs via Homebrew (brew tap synseqack/aict && brew install aict, which also adds shell completions), go install, or building from source, and has exactly one external dependency - the official MCP Go SDK, used only by the mcp subcommand - with every one of the 34 coreutils implemented in pure Go standard library. It's strictly read-only, makes no network requests (MIME detection uses Go's stdlib, not an HTTP lookup) and collects no telemetry, so it's safe to run in a sandboxed agent environment.
Who it's for
Teams building or running coding agents that repeatedly shell out to ls, grep, find, diff, and similar tools and want structured, unambiguous output the agent can consume directly - accepting a real token and latency cost in exchange for fewer round-trips and, per the project's own benchmark, fewer wrong answers from misparsed plaintext. aict is released under the MIT license.
Source README
Unix coreutils with XML/JSON output - built for AI agents, not humans.
Install · Quick start · All tools · MCP server · Claude Code · Token cost · Benchmarks · Contributing
The problem
AI agents run ls, grep, and cat and get back human-readable plaintext. Then they spend tokens parsing column positions, guessing field widths, and handling inconsistent formats. This is fragile and wasteful.
-rw-r--r-- 1 user staff 2048 Apr 6 10:00 main.go ← which column is size?
-rw-r--r-- 1 user staff 1024 Apr 6 10:00 utils.go ← what's the language?
drwxr-xr-x 5 user staff 160 Apr 6 10:00 internal ← is this a directory?
The solution
aict reimplements 33 Unix tools with structured output the agent can read directly - no parsing required.
$ aict ls src/
<ls timestamp="1746123456" total_entries="3">
<file name="main.go" path="src/main.go" absolute="/project/src/main.go"
size_bytes="2048" size_human="2.0K" language="go" mime="text/x-go"
binary="false" executable="false" modified="1746120000" modified_ago_s="3456"/>
<file name="utils.go" path="src/utils.go" absolute="/project/src/utils.go"
size_bytes="1024" size_human="1.0K" language="go" mime="text/x-go"
binary="false" executable="false" modified="1746120000" modified_ago_s="3456"/>
<directory name="internal" path="src/internal" modified="1746120000"/>
</ls>
Every field is labeled. Paths are always absolute. Timestamps are Unix integers. Language and MIME type are detected automatically - zero parsing needed.
Install
Homebrew (macOS)
brew tap synseqack/aict
brew install aict
This installs aict plus shell completions for bash and zsh. The MCP server is built in: aict mcp.
Go Install
go install github.com/synseqack/aict@latest
Build from Source
git clone https://github.com/synseqack/aict
cd aict
go build -o aict .
Verify install:
aict --helpshould list all available tools.
Quick start
# Default: XML output (best for AI agents)
aict ls src/
aict grep "func" . -r
aict cat main.go
aict diff old.go new.go
# JSON output
aict ls src/ --json
# Plain text (same as the original Unix tools)
aict ls src/ --plain
# Enable XML globally for all aict calls
export AICT_XML=1
Tools
34 tools across 6 categories. Every tool supports --xml (default), --json, and --plain. Compact mode (default) uses short attribute names to save tokens - use --no-compact for verbose output or --dict to see the mapping.
| Category | Tools |
|---|---|
| File inspection | cat head tail file stat wc |
| Search & compare | ls find grep diff |
| Path utilities | realpath basename dirname pwd |
| Text processing | sort uniq cut tr sed awk |
| Data & archives | jq tar |
| System & environment | env system ps df du checksums md5sum sha1sum sha256sum |
Additional: git (status, diff, log, ls-files, blame) · completions (bash/zsh/fish) · doctor (self-diagnostic)
Output format
All tools follow the same conventions:
| Field | Compact (default) | Verbose (--no-compact) |
|---|---|---|
| Paths | p, a |
path, absolute |
| Timestamps | t (epoch int) + ma (_ago_s) |
timestamp + modified_ago_s |
| Sizes | s (bytes) + sh (size_human) |
size_bytes + size_human |
| Booleans | 1 / 0 |
"true" / "false" |
| Errors | <error c="" msg=""/> |
<error code="" msg=""/> |
| Empty results | Valid XML with zero counts, never an error | Same |
MCP server
aict mcp exposes all tools as callable MCP functions via stdio transport. AI assistants call them natively - no shell wrapping needed. The MCP server is a subcommand of the main binary.
Configure Claude Desktop (~/.config/claude/claude_desktop_config.json):
{
"mcpServers": {
"aict": {
"command": "aict",
"args": ["mcp"]
}
}
}
If aict is not in PATH, use its full path:
{
"mcpServers": {
"aict": {
"command": "/usr/local/bin/aict",
"args": ["mcp"]
}
}
}
Claude Code integration
Add to ~/.claude.json:
{
"mcpServers": {
"aict": {
"command": "aict",
"args": ["mcp"]
}
}
}
Once connected, Claude Code can call ls, grep, diff, and all other tools as native functions with typed arguments and structured JSON results.
Token cost
Honest numbers first: aict output costs 1.1-7.8× more tokens per task than terse GNU output (measured with tiktoken o200k_base). What those tokens buy: fewer round-trips and zero parsing ambiguity. Where plaintext needs 2-4 chained calls (ls then file for languages, find then stat per hit), aict answers in one.
| Task | GNU | aict | Tokens (GNU → aict) |
|---|---|---|---|
| List dir with size/type/language | 2 calls | 1 call | 246 → 820 · 3.3× |
| Read file + line count + type | 3 calls | 1 call | 173 → 274 · 1.6× |
Find .go files with size/mtime |
4 calls | 1 call | 47 → 367 · 7.8× |
| Grep with file/line/context | 1 call | 1 call | 141 → 326 · 2.3× |
| Diff with change types | 1 call | 1 call | 167 → 192 · 1.2× |
Every extra plaintext call is a full agent turn - model inference, tool-call overhead, and intermediate output all land in the context window anyway, none of which the token counts above include. And in this very benchmark, file(1) misidentified a Go source file as "C source"; aict labeled it go.
In a live agent eval (opencode, same task 3× per toolchain), the aict-equipped agent generated ~46% fewer output tokens (median 265 vs 487) and was correct 3/3 - the GNU-equipped agent shipped a flawed report in the run where it trusted file(1)'s language detection. See benchmarks/TOKENS.md for both methodologies; reproduce with go run ./cmd/tokenbench.
Use --plain when you only need raw content.
Benchmarks
aict trades some speed for semantic richness (language detection, MIME typing, absolute paths). The overhead is intentional. Startup cost is ~3.6 ms per invocation.
| Tool | GNU | --plain |
--xml |
Notes |
|---|---|---|---|---|
diff (1000 lines) |
0.9 ms | 1.9 ms · 2.1× | 2.1 ms · 2.4× | ✅ Myers O(ND) |
wc (100k lines) |
6.1 ms | 16 ms · 2.6× | 17 ms · 2.7× | ✅ |
awk (10k lines) |
4.1 ms | 12 ms · 2.9× | 11 ms · 2.6× | ✅ |
sed (10k lines) |
3.3 ms | 14 ms · 4.2× | 16 ms · 4.9× | ✅ |
find (deep tree) |
1.9 ms | 13 ms · 6.8× | 15 ms · 8.0× | ✅ |
ls (1000 files) |
4.0 ms | 51 ms · 12.9× | 70 ms · 17.7× | MIME+lang detection per file |
cat (100k lines) |
1.4 ms | 24 ms · 16.4× | 31 ms · 21.6× | line-by-line scan + encoding detect |
grep (100k lines) |
1.3 ms | 119 ms · 88× | 130 ms · 96× | Go regexp vs GNU SIMD |
Medians from 5 runs on Linux/amd64. See benchmarks/ for methodology and make bench to reproduce.
Use --plain to skip enrichment when you only need raw content.
FAQ
Why XML and not JSON by default?
XML attributes are denser in a context window. <file size="1024" lang="go"/> is shorter than {"size":1024,"lang":"go"}. Use --json if you prefer JSON - the structure is identical.
Why not pipe GNU tools to jq?
ls, cat, stat, find, diff, and wc don't output JSON. jq can't help with them. aict provides structured output for the entire toolchain, not just grep. (aict also ships its own jq for querying JSON files with path expressions.)
How does this compare to ripgrep?
ripgrep is much faster for pure search. aict grep adds language detection, MIME type, and a consistent output format shared with every other tool. Use ripgrep for speed-critical search; use aict when the agent needs structured context.
How does this compare to eza / lsd?
eza and lsd are better ls for humans - great colors and formatting. aict outputs data structures, not formatted tables. They're solving different problems.
Does it work on Windows?
ls, cat, stat, wc, find, diff, grep, head, tail, sort, uniq, cut, tr, sed, awk, jq, tar, checksums, df, and path utilities work on Windows. system is Linux/macOS; ps is Linux-only (it reads /proc). Unsupported platforms get a structured <error> element, never a crash.
Is this safe to run in a sandboxed environment?
Yes. aict is strictly read-only. No network requests (MIME detection uses the Go stdlib, not HTTP). No telemetry. No data collection. It only reads paths you explicitly pass to it.
How many dependencies does it have?
One: the official MCP Go SDK, used only by the aict mcp subcommand. All 33 tools and every internal package are pure Go standard library - enforced as a hard constraint in AGENTS.md.
FAQ
Common questions
Discussion
Questions & comments · 0
Sign In Sign in to leave a comment.