Every integration you will ever build eventually funnels through email. Ticketing systems, invoice ingestion, newsletter archives, mailbox sync, test-report pipelines — all of them have to take a raw .eml produced by some mail client written in 2003 and turn it into a subject line, a body and a list of attachments. That conversion is where production breaks: encoded words, nested multipart/mixed inside multipart/alternative, attachments that claim to be text/plain, charsets that lie, and 40 MB messages that your parser happily loads into memory.

This article compares the four MIME parsers that actually survive real traffic in 2026 — mailparser (Node), enmime (Go), Apache Mime4j (JVM) and mail-parser (Rust) — with live GitHub figures, working code, and the specific failure modes each one handles well or badly.

TL;DR: The Verdict

Use mailparser in Node services — it has the politest API in the field and normalises text, html and CID-referenced inline images for you. Use enmime in Go — same ergonomics, plus a Errors slice that tells you which parts were malformed instead of throwing. Use Apache Mime4j on the JVM when you need streaming over a 100,000-message mailbox rather than a DOM in memory. Use mail-parser in Rust when throughput matters — it is the fastest of the four and the only one with a genuinely compact, allocation-light design.

The Contenders At A Glance

Live GitHub figures from 2026-10-09/10:

PropertymailparserenmimeApache Mime4jmail-parser
Repositorynodemailer/mailparserjhillyerd/enmimeapache/james-mime4jStalwartLabs/mail-parser
Stars1,670 ⭐523 ⭐67 ⭐ (mirror)461 ⭐
Last push2026-10-082026-10-012026-09-142026-10-08
Language / runtimeJavaScript / NodeGoJava / JVMRust
Parse modelDOM (whole message)DOM (whole message)DOM or streaming event handlerDOM + zero-copy borrows
RFC 2047 header decodingYes, automaticYes, GetHeader()Yes, DecodeMonitor configurableYes
Inline CID imagesattachments[].cid + textAsHtmlInlines slicemanual (Multipart walk)text_bodies/html_bodies + parts
Malformed-input policyLenient, best-effortLenient, records ErrorsConfigurable (PERMISSIVE, STRICT)Lenient with repair heuristics
Attachment streamingNoNoYesPartial
LicenseMITMITApache-2.0Apache-2.0 / MIT

And the ten-second decision table:

Your taskPickWhy
Parse inbound webhook payloads in a Node APImailparsersimpleParser() gives you subject, text, html and attachments in one await
Build a Go mail gateway or ingestion workerenmimeDecoded headers, Errors instead of exceptions, stable API since 2018
Index a 500 GB mail archive on the JVMApache Mime4jMimeStreamParser never materialises the whole message
Parse 1M messages in a Rust batch jobmail-parserFastest parser here; designed for throughput, not ergonomics
Do it with zero dependencies in Pythonstdlib email + policy.defaultModern policy handles RFC 2047 and address headers correctly

mailparser — The Node Default

mailparser is maintained by the Nodemailer project, which is the same team most Node developers already trust for outbound mail. That symmetry — Nodemailer out, mailparser in — is why it became the default.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
const fs = require('fs');
const { simpleParser } = require('mailparser');

const raw = fs.readFileSync('invoice.eml');
const mail = await simpleParser(raw, {
  // 20 MB attachment ceiling; anything larger is dropped with a size flag
  maxHtmlLengthToParse: 10 * 1024 * 1024,
});

console.log(mail.subject);            // already RFC 2047-decoded
console.log(mail.from.value[0].address);
console.log(mail.text);               // plain-text alternative
console.log(mail.html);               // HTML alternative
console.log(mail.headers.get('message-id'));

for (const att of mail.attachments) {
  console.log(att.filename, att.contentType, att.size, att.cid ?? '-');
  fs.writeFileSync(`/tmp/${att.filename}`, att.content);
}

Two things make it pleasant: mail.from is a parsed address object rather than a raw string, and mail.textAsHtml converts the plain-text body into HTML with inline cid: references rewritten to data: URLs — which is exactly what you need to render a message in a browser without a separate attachment server.

Watch out: simpleParser reads the whole message into a Buffer and decodes every attachment. On a mailbox sync job, that is your memory ceiling, not your CPU.

enmime — MIME Handling for Go

enmime (523 ⭐) is the Go ecosystem’s mature answer, and its distinguishing feature is that it does not throw on malformed input. It returns an envelope plus a slice of errors describing which MIME parts failed, so a single broken Content-Type header does not abort a batch.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
import (
	"github.com/jhillyerd/enmime"
	"os"
)

f, _ := os.Open("message.eml")
defer f.Close()

env, err := enmime.ReadEnvelope(f)
if err != nil {
	return err
}

subject := env.GetHeader("Subject")   // RFC 2047 decoded
text, html := env.Text, env.HTML

for _, att := range env.Attachments {
	os.WriteFile("/tmp/"+att.FileName, att.Content, 0o600)
}

for _, inl := range env.Inlines {
	// inline images referenced by Content-ID from the HTML body
	fmt.Println(inl.FileName, inl.Header.Get("Content-ID"))
}

if len(env.Errors) > 0 {
	log.Printf("partially malformed message: %v", env.Errors)
}

The Inlines / Attachments split is the detail that saves you work: inline images are separated from real attachments, so a “download all attachments” endpoint does not accidentally hand users a logo.

Apache Mime4j — The JVM Workhorse You Never Notice

Mime4j is the parsing engine underneath Apache James, the enterprise mail server, so it has seen more hostile input than any other library in this list. Its star count (67, and only on a mirror) understates its footprint entirely — it is a 15-year-old Apache project built for mailboxes, not demos.

For a single message, the DOM API is straightforward:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
import org.apache.james.mime4j.message.DefaultMessageBuilder;
import org.apache.james.mime4j.dom.Message;
import org.apache.james.mime4j.dom.Multipart;
import org.apache.james.mime4j.dom.BodyPart;

MessageBuilder builder = new DefaultMessageBuilder();
builder.setMimeEntityConfig(MimeConfig.PERMISSIVE);   // real-world tolerant mode

try (InputStream in = new FileInputStream("message.eml")) {
    Message msg = builder.parseMessage(in);
    System.out.println(msg.getSubject());
    for (BodyPart part : ((Multipart) msg.getBody()).getBodyParts()) {
        System.out.println(part.getMimeType() + " " + part.getFilename());
    }
}

The reason to choose Mime4j is the streaming API: MimeStreamParser pushes events to your handlers (startMessage, body, startMultipart, field) as bytes arrive. Parse a 100 million-message archive with a constant memory footprint, and configure MimeConfig for strict or permissive parsing per source. No other parser in this table offers that combination.

mail-parser — Rust Throughput

StalwartLabs/mail-parser (461 ⭐, pushed 2026-10-08) is a newer entrant from the Stalwart mail-server team, written with the explicit goal of parsing millions of messages without an allocation storm.

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
use mail_parser::MessageParser;

let raw = std::fs::read("message.eml").unwrap();
let message = MessageParser::default()
    .parse(&raw)
    .expect("no message");

println!("{:?}", message.subject());
println!("{:?}", message.date());
println!("{:?}", message.text_body(0));
println!("{:?}", message.html_body(0));

for part in message.attachments() {
    println!("{:?} {:?}", part.attachment_name(), part.content_type());
}

The library exposes explicit part accessors (text_body(index), html_body(index)) instead of collapsing everything into one field — helpful when a message contains three text/plain parts and you need to know which is which. It is Apache-2.0/MIT dual-licensed, so it drops into either kind of project.

Python, in Three Lines, With No Dependencies

Python’s standard library got good at this in 3.6+ because of the modern email policy. policy.default turns on RFC 2047 header decoding and the address-header registry, which fixes the most common complaint about the old API:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
import email
from email import policy

with open("message.eml", "rb") as fh:
    msg = email.message_from_binary_file(fh, policy=policy.default)

print(msg["Subject"])                     # decoded, spaces collapsed correctly
print(msg["From"].addresses[0].addr_spec)

body = msg.get_body(preferencelist=("plain", "html"))
print(body.get_content())

for part in msg.iter_attachments():
    name = part.get_filename()
    data = part.get_payload(decode=True)  # transfer-encoding decoded
    print(name, part.get_content_type(), len(data or b""))

If you are on Python and writing your own header regexes, stop. The stdlib path handles encoded words, folded headers and address lists correctly, and — unlike a regex — it will not be defeated by a header value containing a colon.

Pitfalls: Where MIME Parsers Bite in Production

  • Never trust Content-Type. Inline images labelled application/octet-stream, HTML bodies labelled text/plain, and multipart boundaries that appear inside quoted text are all routine. Decide content types from sniffed bytes and declared intent, not one header.
  • Cap attachment sizes before decoding. Even a DOM parser will decode a 2 GB attachment into memory if you let it. Set a per-attachment and per-message limit, and record the drop (enmime surfaces this in Errors).
  • Decode transfer encodings with a hard limit on expansion. A small quoted-printable or base64 payload can expand enormously; treat the declared size as untrusted.
  • Sanitise HTML before rendering. Parsing is not sanitisation. A parsed html field is attacker-controlled markup — run it through an allow-list sanitiser before it reaches a browser.
  • Handle charset fallbacks explicitly. When a message declares charset=iso-8859-1 but actually contains UTF-8, decoding with errors="replace" loses data silently. Try declared, then UTF-8, then latin-1, and record which one won.
  • Thread on Message-ID/References, not on subject. Subject lines are rewritten by every mailing list on the planet; In-Reply-To is not.
  • Parse Date with a tolerant parser. Two-digit years, missing timezone offsets and obsolete zone names (EST without a numeric offset) are all legal in the wild.
  • Keep the raw bytes. Store the original .eml alongside your parsed representation. Every parser fixes bugs, and re-parsing beats re-requesting mail from a provider that has already purged it.

If you are building the server side of this pipeline rather than the parser, our comparison of self-hosted SMTP relays covers the delivery layer, and the self-hosted email archiving guide covers long-term storage. For attachment policy and legal footers — another place where MIME structure matters — see the email disclaimer management comparison.

FAQ

Which MIME parser should I use in Node.js? mailparser from the Nodemailer project. simpleParser() returns decoded headers, plain-text and HTML bodies, parsed address objects and an attachment array with cid values in a single await, which removes almost all of the boilerplate that makes MIME handling painful.

How do I parse emails without loading the entire message into memory? Use Apache Mime4j’s MimeStreamParser on the JVM, or mail-parser in Rust with its borrow-based design. DOM parsers like mailparser and enmime load the full message, which is fine for individual requests but becomes the bottleneck when indexing a large archive.

Why do my email subjects look like =?UTF-8?B?...?=? That is RFC 2047 encoded-word syntax, used for non-ASCII headers. Every parser in this comparison decodes it automatically — mailparser, enmime via GetHeader(), Mime4j, mail-parser, and Python’s email with policy.default. If you see raw encoded words, you are reading a header with a legacy API or a regex.

How should I handle inline images in HTML emails? Look for attachments whose Content-ID is referenced from the HTML body as cid:.... mailparser does this for you (attachments[].cid, plus textAsHtml rewriting), and enmime separates them into an Inlines slice so they never get mixed with real attachments.

Should I store the raw .eml file or just the parsed JSON? Store both. The parsed representation is what your application queries, but the raw message is the authoritative record — it lets you re-parse when you upgrade a library, prove what was actually received during a dispute, and recover fields your parser dropped.

Is parsing an email enough to render it safely in a browser? No. Parsing validates structure, not safety. Treat the HTML body as untrusted input and run it through an allow-list sanitiser with scripts, event handlers and remote loaders removed before it ever reaches a browser context.

Can I parse MIME in Python without third-party packages? Yes. email.message_from_binary_file(fh, policy=policy.default) gives you RFC 2047 header decoding, structured address headers, get_body() for the preferred alternative, and iter_attachments() with transfer-encoding already decoded.


💰 想测试你的市场判断力?我用 Polymarket 做预测市场交易——这是全球最大的预测市场平台,从大选结果到技术监管时间线,什么都可以押注。和赌博不同,这是真正的信息市场:你懂的信息越多,胜率越高。我靠预测技术相关事件的走向已经赚了不少。用我的邀请链接注册:Polymarket.com