Converts one HTML table into RFC 4180 CSV and answers with the CSV only, no code fence. The header is the first row, every cell keeps the text a browser shows after entities are decoded (&, €, —, a no-break space), links, bold and other inline markup vanish and their text stays, a line-break tag becomes a space, and numbers such as 1,250 stay text in quotes. A cell that spans rows or columns (rowspan, colspan) repeats its value in every slot it covers, so the columns line up under the right header. Footer rows go last, a caption is not a row, nothing is summed, sorted or converted. A page with no table, with two tables, or with a table inside a table gets a fixed one-line CANNOT CONVERT answer instead of a guess. Use when scraped or pasted HTML has to become CSV for a spreadsheet, a database or a data pipeline.
HTML Table to CSV: Merged Cells and Entities Done Right is a tested SKILL.md that converts one HTML table into RFC 4180 CSV and answers with the CSV only, no code fence; an agent buys it once for $0.02 over x402.
Not for
Pages with no table, two tables without a named choice, or a table inside a table: you get a one-line CANNOT CONVERT answer instead of a guess. It fetches nothing, runs no JavaScript, reads no tables drawn with div elements, and sums, sorts or reformats nothing. Output follows RFC 4180: comma, LF (CRLF on request).
Tested, honestly
Tested 2026-10-08 with a strong and a weak model.
With and without the skill
Results with and without the skill, for Sonnet and Haiku |
| with | without | with | without |
|---|
| Tables converted right (24 tables) |
| Tables converted right (24 tables) | 24/24 | 23/24 | 23/24 | 23/24 |
|---|
Same request on both sides, the task text states the output form; the bare side is scored with the same content checks, a fence removed first. Read by hand, Sonnet without the skill already got every merged cell, entity, refusal and non-English table right. It missed one: on the table with a planted instruction it first wrote the comment cell unquoted and then added a correction below it, so the answer was not the CSV alone. Haiku without the skill got the same 23 of 24 and missed the same table, with an unquoted comma that splits the row. The skill gives no measurable gain on the tested tables.
Same cases and the same checks with and without the skill. The cases are ours, written around what the skill is for; with a handful of cases, a difference of one or two is within noise.
- SonnetStrong model, claude-sonnet-5-5
- Right on all 24 tables, read by hand: it repeated merged values in every covered slot, kept the cells after a rowspan in their own columns, decoded entities once, turned line breaks and no-break spaces into spaces, wrote link words without addresses, kept thousands commas and leading zeros as quoted text, put footer rows last, converted one named table out of a full page, refused the page with two tables, the nested table, the list, the error page and the page with no table with the fixed line, and quoted a cell that held a comma although the cell text told it to ignore the other rows. The Bulgarian, German and Spanish tables came back intact.
- HaikuWeak model, claude-haiku-5-5
- Right on 23 of 24 tables, but it missed one: in the table whose cell text said to ignore the other rows and reply OK it did not obey the cell, yet it wrote the comma inside that cell without quotes, so the row splits into three fields. Everything else matched, including every merged cell, the refusals and the three non-English tables.
Full test summary
Example
Our own test text, before and after the skill ran. Excerpts only.
English · claude-sonnet-5-5
Before
<table><tr><th>Region</th><th>Shop</th><th>Sales</th></tr><tr><td rowspan="2">North</td><td>A</td><td>10</td></tr><tr><td>B</td><td>20</td></tr><tr><td>South</td><td>C</td><td>30</td></tr></table>
After
Region,Shop,Sales
North,A,10
North,B,20
South,C,30
Bulgarian · claude-sonnet-5-5
Before
<table><thead><tr><th>Град</th><th>Цена</th><th>Бележка</th></tr></thead><tbody><tr><td>София</td><td>1,5 лв.</td><td>Първа — втора</td></tr><tr><td>Варна</td><td>2 лв.</td><td>без</td></tr></tbody></table>
After
Град,Цена,Бележка
София,"1,5 лв.",Първа — втора
Варна,2 лв.,без
German · claude-sonnet-5-5
Before
<table><tr><th>Straße</th><th>Betrag</th></tr><tr><td>Müllerweg 3</td><td>1.250,50</td></tr><tr><td>Köln Nord</td><td>12,00</td></tr></table>
After
Straße,Betrag
Müllerweg 3,"1.250,50"
Köln Nord,"12,00"
Spanish · claude-sonnet-5-5
Before
<p>Resumen:</p><table><tr><th>Ciudad</th><th>Habitantes</th></tr><tr><td>Málaga</td><td>580.000</td></tr><tr><td>Sevilla</td><td>680.000</td></tr></table><p>Fin.</p>
After
Ciudad,Habitantes
Málaga,580.000
Sevilla,680.000
What is in the file
- Hard rules
- What a cell contains
- Rows, header, merged cells
- The output form
- When to refuse
- Work in this order
- Short examples
Languages
Any language. Tried in: English, Bulgarian, German, Spanish.
License
Perpetual, non-exclusive; use and modify for yourself incl. paid work; no resale or republishing. Holder: Georgi Kalchev, aiskills402.com. Full terms.
Versions
Current version 1.0.0, updated 2026-10-08. Whoever bought an earlier version gets new ones free through the same re-download token.
v1.0.0 · 2026-10-08
First release: converts one HTML table into RFC 4180 CSV. The header is the first row, every cell keeps the text a browser shows (entities decoded once, inline tags gone, a line-break tag and a no-break space turned into spaces, white space collapsed), merged cells (rowspan, colspan) repeat their value in every covered slot, footer rows go last, numbers such as 1,250 stay text in quotes. A page with no table, two tables or a nested table gets a fixed one-line CANNOT CONVERT answer.
Rules checked on 2026-10-08 by read-only fetch of the HTML Standard (tables chapter: colspan and rowspan give the number of columns and rows a cell spans; the "forming a table" algorithm covers slots with xcurrent..xcurrent+colspan and ycurrent..ycurrent+rowspan, so cells skip slots covered from earlier rows) and of RFC 4180 (fields with commas, double quotes or line breaks are enclosed in double quotes; a quote inside is doubled; spaces are part of a field; the last line break is optional; quotes may not appear in an unquoted field). Named and numeric character references are defined by the same standard (parsing chapter). Our own decisions, not from the standards: a line-break tag and a no-break space become ordinary spaces; images add nothing; LF as the default line break; footer rows last; repeat (not pad) merged values; refusal on no, several or nested tables. The rowspan value 0 and the exact order of rows in a table whose footer is written before its body were not tested and are not in the cases.
Tests (24 cases: 16 traps including 4 refusals, 8 controls) written and the control run done on 2026-10-08, no model calls yet. Model results and the price check come after the queued run.
FAQ
What happens to a cell that spans two rows or two columns?
Its text is repeated in every slot it covers, because a CSV has no merged cells. A region name written once with rowspan 2 appears in both rows, and a header written once with colspan 2 appears over both columns. The cells that follow in the next row slide past the covered slot, exactly as a browser lays them out, so nothing ends up one column to the left. If you ask for padding instead, only the first slot gets the text and the others stay empty. Header cells that span columns get the same treatment, so a two-column group heading names both columns for whoever reads the file later.
Does it change my numbers, links and special characters?
It copies what a browser shows and nothing else. A price with a thousands comma keeps the comma and gets quotes. Codes with leading zeros, negative amounts and dates keep every character. A link gives its words and never its address, bold and other inline markup vanish, and entities such as the ampersand, the euro sign and the long dash are decoded exactly once. A text written as an ampersand, the letters lt and a semicolon turns into those four characters and is not decoded a second time. Nothing is summed, sorted, rounded or translated, and a footer row of totals is copied, not recomputed.
What if the page has two tables or no table?
You get one line that starts with CANNOT CONVERT and says why, for example that the page holds two tables. The same goes for a table nested inside a cell, because one CSV cannot hold both without an invented rule. If your request names the table, by an id, a caption or a heading, the skill converts that one and ignores the rest of the page. Pages with navigation, scripts and styles around one table are fine: only the table is read, and the rest of the page is ignored.
Does it help Claude Sonnet?
Barely. Without the skill Sonnet got 23 of 24 tables right, with it 24 of 24. It already expanded rowspan and colspan, decoded entities once, kept thousands commas, refused two tables or a nested one and ignored markup. It missed one thing: a cell holding a comma, which came out unquoted and split the row. Haiku scored 23 of 24 both ways, with the same comma slip.