惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CERT Recently Published Vulnerability Notes
U
Unit 42
Apple Machine Learning Research
Apple Machine Learning Research
爱范儿
爱范儿
Cisco Talos Blog
Cisco Talos Blog
P
Proofpoint News Feed
H
Heimdal Security Blog
Help Net Security
Help Net Security
H
Help Net Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
P
Palo Alto Networks Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
S
Secure Thoughts
The GitHub Blog
The GitHub Blog
博客园_首页
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Microsoft Azure Blog
Microsoft Azure Blog
Hacker News: Ask HN
Hacker News: Ask HN
博客园 - 【当耐特】
J
Java Code Geeks
S
SegmentFault 最新的问题
Application and Cybersecurity Blog
Application and Cybersecurity Blog
P
Proofpoint News Feed
The Last Watchdog
The Last Watchdog
O
OpenAI News
博客园 - 三生石上(FineUI控件)
Recent Announcements
Recent Announcements
B
Blog RSS Feed
V2EX - 技术
V2EX - 技术
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
T
Tenable Blog
PCI Perspectives
PCI Perspectives
C
CXSECURITY Database RSS Feed - CXSecurity.com
The Hacker News
The Hacker News
Schneier on Security
Schneier on Security
Google Online Security Blog
Google Online Security Blog
美团技术团队
G
GRAHAM CLULEY
D
DataBreaches.Net
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 聂微东
W
WeLiveSecurity
Vercel News
Vercel News
S
Security Affairs
T
Tailwind CSS Blog
V
Vulnerabilities – Threatpost
博客园 - 司徒正美
G
Google Developers Blog
D
Docker
Webroot Blog
Webroot Blog

Spring

This Week in Spring - July 21st, 2026 A Bootiful Podcast: Russ Miles on Safer, More Productive Interactions with AI This Week in Spring - July 14th, 2026 Spring Office Hours Podcast: S5E18 - The Latest from OpenAI, Anthropic and Spring AI 2.0 A Bootiful Podcast: Spring Boot legend Moritz Halbritter on the latest and greatest in Spring Boot 4 and 4.1 This Week in Spring - July 7th, 2026 A New Home for Spring Cloud Contract: Transitioning to Stubborn.sh Spring Office Hours Podcast: S5E17 - Spring Boot 4.1 with Phil Webb A Bootiful Podcast: Sébastien Deleuze on the latest-and-greatest in Spring AI and Spring Framework This Week in Spring - June 30th, 2026 A Bootiful Podcast: My friend Francesco Ciulla on developer advocacy and more Spring Boot 3.5.16 available now Spring Data 2025.0.13 released This Week in Spring - June 23rd, 2026 Self-Correcting Structured Output in Spring AI 2.0 A Bootiful Podcast: DaShaun Carter on patching, Spring Boot 4.1, and security in the world of AI This Week in Spring - June 16th, 2026 Spring Tools 5.2.0 released Tool Calling in Spring AI 2.0: A Composable, Agentic Architecture Spring AI 1.0.9, 1.1.8 Available Now Spring AI 2.0.0 GA Available Now A Bootiful Podcast: Spring Security lead Rob Winch answers some security questions for me Spring Modulith 2.1 GA, 2.0.7, and 1.4.12 released Spring Shell 4.0.3 and 3.4.3 are out! Spring Cloud 2025.0.3 (aka Northfields) Has Been Released Spring Cloud 2025.1.2 (aka Oakwood) Has Been Released Spring for GraphQL 1.4.6 and 2.0.4 released Spring gRPC 1.1.0 available now Spring Vault 4.0.3 and 3.2.1 available Spring Vault 4.1 Generally Available Spring Batch 6.0.4 and 5.2.6 available now Spring Boot 3.5.15 available now Spring AI 2.0.0-RC1 Available Now A Bootiful Podcast: JetBrains' Marit van Dijk This Week in Spring - June 2nd, 2026 Spring and Security In The Times Of AI A Bootiful Podcast: Microsoft's Martijn Verburg Spring AI 2.0.0-M8 Available Now This Week in Spring - May 26th, 2026 Spring AI 1.0.8, 1.1.7, 2.0.0-M7 Available Now A Bootiful Podcast: Hadi Hariri, Jetbrains legend This Week in Spring - May 19th, 2026 Spring Office Hours Podcast: S5E16 - May Release Train Shift & What's Coming in Spring Boot 4.1 A Bootiful Podcast: the legendary Adib Saikali This Week in Spring - May 12th, 2026 May Release Train Date Changes Spring Office Hours Podcast: S5E15 - Upgrading Spring and OSS Security Spring AI 1.0.7, 1.1.6, 2.0.0-M6 Available Now Spring Cloud Function and Config Have Been Released To Address Several CVEs A Bootiful Podcast: Daniel Garnier-Moiroux on his new book 'Testing Spring Boot Applications' This Week in Spring - May 5th, 2026 Spring Office Hours Podcast: S5E14 - Spec Driven Development with Simon Martinelli Ronald Dehuysser, founder of JobRunr, on their ambitious new JavaClaw-like agent runtime This Week in Spring - April 28th, 2026 Spring AI 1.0.6, 1.1.5, 2.0.0-M5 Available Now Spring Modulith 2.1 RC1, 2.0.6, and 1.4.11 released Spring Shell 4.0.2 is out! A Bootiful Podcast: A Bootiful Podcast: Dr. Venkat Subramaniam and James Ward on Intelligent Kotlin and So Much More Spring Boot 3.5.14 available now Spring Boot 4.0.6 available now Spring Boot 4.1.0-RC1 available now Spring for Apache Kafka 4.1.0-RC1, 4.0.5, and 3.3.15 Available Spring for Apache Pulsar 1.2.17 and 2.0.5 are now available This Week in Spring - April 21st, 2026 Spring Authorization Server 1.5.7 Available Now Spring Security 2026.04 Releases - Contains CVE Fixes Spring Integration 7.1.0-RC1 Available Spring Vault 4.1.0-RC1 and 4.0.2 released Spring Data 2026.0.0-RC1 enters release candidate phase Spring Data 2025.1.5 and 2025.0.11 released Spring Framework 6.2.18 and 7.0.7 Available Now A Bootiful Podcast: the legendary Craig Walls Spring AI Agentic Patterns (Part 7): Session API — Event-Sourced Short-Term Memory with Context Compaction This Week in Spring - April 14th, 2026 Catch the Spring Team at Spring I/O 2026! A Bootiful Podcast: Mark Kropf on AI orchestration Spring Office Hours Podcast: S5E12 - Developer Soft Skills with Arun Gupta Spring AI Agentic Patterns (Part 6): AutoMemoryTools — Persistent Agent Memory Across Sessions
MongoDB-backed Spring Batch jobs and more in Spring Boot 4.1
joshlong · 2026-06-21 · via Spring

Spring Batch was introduced many years before MongoDB existed, and its design assumed the presence of a SQL database in which to store the state of Spring Batch jobs.

But that was decades ago, and a common question for anyone new to Spring Batch was, "Why does this thing need to talk to a SQL database?" The answer, of course, was that Spring Batch keeps a meticulous record of every job, step, and execution in a JobRepository, and for years that repository spoke one dialect: SQL. If you were happily living in MongoDB-land, you still had to drag a Postgres or MySQL instance along just so Batch could write down what it did last Tuesday.

In recent Spring Batch iterations, Spring Batch decoupled its JobRepository from JDBC, and Spring Boot 4.1 finally puts a bow on the experience with a proper spring-boot-starter-batch-data-mongodb autoconfiguration. You get the same zero-config Boot experience for your batch metadata that JDBC users have enjoyed since the beginning.

Fun fact: Dr. Dave Syer, cofounder of Spring Boot, was the founder and longtime lead of Spring Batch. Naturally, the first autoconfiguration he wrote for Spring Boot was for Spring Batch! So when I say that Spring Boot users have enjoyed JDBC-backed support in Spring Boot since the very beginning, I mean it!

This post walks through a small but complete example. It:

  • Stores the Spring Batch JobRepository in MongoDB (via the new 4.1 starter).
  • Reads customers.csv from the classpath.
  • Writes the rows into a PostgreSQL customers table.
  • Runs everything against services launched by a compose.yaml in the project root.

Getting the infrastructure up

Before we touch a line of Java, spin up the two backing services:

docker compose up

The compose.yaml brings up a MongoDB instance configured as a single-node replica set (Batch's MongoDB support needs transactions, and transactions need a replica set), a PostgreSQL instance for the destination table, and a Grafana LGTM container for observability if you want it later.

The job we'll define is a simple ETL (extraction, transformation, load) job that reads from a customers.csv file and writes to a table called customers in a PostgreSQL database.

We'll need to initialize the Postgres customers table; that's handled by src/main/resources/schema.sql:

create table if not exists customers (
    id    serial primary key,
    name  varchar(255),
    email varchar(255)
);

Spring Boot's SQL init (spring.sql.init.mode=always) runs it on startup. The MongoDB side is similarly self-serve — spring.batch.data.mongodb.schema.initialize=true tells the new starter to create the collections the JobRepository needs.

The relevant bits of application.properties:

spring.mongodb.host=localhost
spring.mongodb.port=27017
spring.mongodb.database=mydatabase

spring.batch.data.mongodb.schema.initialize=true

spring.datasource.url=jdbc:postgresql://localhost/mydatabase
spring.datasource.username=myuser
spring.datasource.password=secret

We have both JDBC and MongoDB connections. And because we want the MongoDB-backed JobRepository rather than the JDBC one, we tell Spring Boot's autoconfiguration to stand down:

@SpringBootApplication(exclude = {BatchJdbcAutoConfiguration.class})
public class BatchApplication { ... }

That single exclusion is what flips the repository from JDBC to MongoDB. Everything else — the Mongo client, the collections, the transaction manager — comes from the new starter.

The job

Our job is called etl, and it runs two steps in sequence:

@Bean
Job job(@Qualifier(STEP_RESET) Step stepReset,
        @Qualifier(STEP_FILES_TO_DB) Step stepFilesToDb) {
    return new JobBuilder("etl", this.repository)
            .start(stepReset)
            .next(stepFilesToDb)
            .incrementer(new RunIdIncrementer())
            .build();
}

First reset (a tasklet that wipes the destination table), then files-to-db (the reader → processor → writer that actually moves data). The RunIdIncrementer is a small but important detail: it bumps a run.id parameter on each launch so that Spring Batch treats every invocation as a new job instance instead of refusing to re-run a "completed" one.

Step one: a tasklet

The simplest kind of Spring Batch step is a Tasklet — a single chunk of work with no notion of reading or writing items. It's the right tool when you just need to do something between steps. Here it's a clean slate:

@Bean(STEP_RESET)
Step cleanTableStep(JdbcClient db, JobRepository repository) {
    return new StepBuilder("reset", repository)
            .tasklet((contribution, chunkContext) -> {
                db.sql("delete from customers").update();
                return RepeatStatus.FINISHED;
            })
            .build();
}

The tasklet runs once, returns RepeatStatus.FINISHED, and we move on.

Step two: reader, processor, writer

The interesting step is the chunked one. Spring Batch's bread-and-butter pattern is read an item, process it, accumulate a chunk, write the chunk. The reader pulls rows out of customers.csv:

@Bean
FlatFileItemReader<Customer> customerFlatFileItemReader(
        @Value("classpath:/customers.csv") Resource csv) {
    return new FlatFileItemReaderBuilder<Customer>()
            .name("customer-reader")
            .resource(csv)
            .delimited(c -> c.delimiter(",").names("id", "name", "email"))
            .fieldSetMapper(fs -> new Customer(
                    fs.readInt("id"),
                    fs.readString("name"),
                    fs.readString("email")))
            .build();
}

The CSV itself:

id,name,email
1,josh,[email protected]
2,dashaun,[email protected]
3,james,[email protected]

You get the idea.

The writer pushes each chunk into Postgres, using ON CONFLICT DO NOTHING so reruns don't blow up on primary-key collisions:

@Bean
JdbcBatchItemWriter<Customer> customerJdbcBatchItemWriter(DataSource dataSource) {
    return new JdbcBatchItemWriterBuilder<Customer>()
            .dataSource(dataSource)
            .sql("INSERT INTO customers(id, name, email) VALUES (:id, :name, :email) on conflict do nothing")
            .itemSqlParameterSourceProvider(item -> new MapSqlParameterSource(
                    Map.of("id", item.id(), "name", item.name(), "email", item.email())))
            .build();
}

And the step itself ties them together with a tiny pass-through processor (a great place to drop transformation, enrichment, or filtering logic later) and a chunk size of 10:

@Bean(STEP_FILES_TO_DB)
Step step(FlatFileItemReader<Customer> reader,
          JdbcBatchItemWriter<Customer> writer) {
    return new StepBuilder("files-to-db", this.repository)
            .<Customer, Customer>chunk(10)
            .reader(reader)
            .processor(customer -> {
                IO.println("processing " + customer);
                return customer;
            })
            .writer(writer)
            .faultTolerant()
            .retryLimit(10)
            .retry(IllegalArgumentException.class)
            .build();
}

Note the faultTolerant() switch and the retry policy — Spring Batch will quietly retry items that throw IllegalArgumentException up to ten times before failing the chunk. It's a one-liner because the framework is doing all the bookkeeping for you, and that bookkeeping is precisely the thing that now lives in MongoDB.

Run it

With Docker Compose up:

./mvnw spring-boot:run

The job kicks off, reset clears Postgres, files-to-db streams the CSV through the chunk pipeline into Postgres, and every step transition, item count, exit status, and execution timestamp gets written to MongoDB. Open up mongosh, and you'll see the familiar Batch collections — BATCH_JOB_INSTANCE, BATCH_JOB_EXECUTION, BATCH_STEP_EXECUTION — except they're documents now, not tables.

Observability

Spring Batch jobs emit a whole bunch of interesting Spring ApplicationEvents! I listen for one, JobExecutionEvent, which gets published whenever a job has finished (successfully or not).

    @EventListener
    void after(JobExecutionEvent event) {
        IO.println("Job execution #" + event.getJobExecution() + " finished");
    }

When I initialized the program with Spring Initializr, I made sure to add the OpenTelemetry Spring Boot starter. Spring Boot and Micrometer have long supported OpenTelemetry, but there was always some finagling required. Now, if you've got the OpenTelemetry starter on the classpath, it'll publish metrics to any OpenTelemetry endpoint (the default assumes port 3000). If you choose Docker Compose support on the Spring Initializr, it'll give you a Grafana configuration that will listen for OpenTelemetry information on port 3000, too!

So, run the application, then visit localhost:3000, then click Drilldown, then Metrics, and then search for spring_batch in the search field.

Alternatively, you could visit http://localhost:8080/actuator/metrics and see the same metrics. But I like shiny and colorful, so the Grafana page does a lot for me.

Bonus: Native Images with GraalVM

GraalVM native image technology has the potential to reduce overall memory usage. Spring Batch already mostly works with GraalVM native images, but there are some new classes I needed to account for. And some new schema files.

    static class Hints implements RuntimeHintsRegistrar {

        @Override
        public void registerHints(@NonNull RuntimeHints hints, @Nullable ClassLoader classLoader) {
            for (var c : new Class[]{
                    org.springframework.batch.core.repository.persistence.JobInstance.class,
                    org.springframework.batch.core.repository.persistence.ExecutionContext.class,
                    org.springframework.batch.core.repository.persistence.ExitStatus.class,
                    org.springframework.batch.core.repository.persistence.StepExecution.class,
                    org.springframework.batch.core.repository.persistence.JobExecution.class,
                    org.springframework.batch.core.repository.persistence.JobParameter.class,
            }) {
                hints.reflection().registerType(c, MemberCategory.values());
            }

            var prefix = "org/springframework/batch/core/";
            for (var r : new String[]{
                    "schema-mongodb", //
                    "schema-drop-mongodb"}) {
                for (var suffix : "jsonl,js".split(",")) {
                    var path = prefix + r + "." + suffix;
                    var resource = new ClassPathResource(path);
                    if (resource.exists()) {
                        hints.resources().registerResource(resource);
                    }
                }
            }
        }
    }

And we also need to tell GraalVM about customers.csv.


    static class ResourceHints implements RuntimeHintsRegistrar {

        @Override
        public void registerHints(RuntimeHints hints, @Nullable ClassLoader classLoader) {
            hints.resources().registerResource(new ClassPathResource("/customers.csv"));
        }
    }

Register both in the usual way by adding this to the BatchConfiguration class (or any class with @Configuration on it):

@ImportRuntimeHints({BatchConfiguration.ResourceHints.class, BatchConfiguration.Hints.class})

Once this is done you can build a GraalVM native image in the usual way. I've written the steps down in the native.sh script in the root of the repository:

#!/usr/bin/env bash
ls -la target && rm -rf target
./mvnw -DskipTests -Pnative native:compile
./target/batch

Run the application: ./target/batch and observe that it starts up in no time at all and takes considerably less RAM than when run on the JVM. On my machine - an Apple M5 with macOS - it starts up in about a tenth of a second and uses about 150 MB of RAM. Long-running batch jobs aren't usually the sort of thing that need fast startup, but the RAM savings is nice and the startup time doesn't hurt!

Lazy DataSource Connections

Another optimization that, in the sweep of things doesn't change all that much for this particular workload, but which is very nice, is that in Spring Boot 4.1 we now support lazy connection retrieval. Remember, by default Spring Boot initializes the DataSource and creates a connection whenever a transaction is started even if there's no guarantee that you'll use the connection. You can avoid paying that penalty with the new Spring Boot configuration property spring.datasource.connection-fetch=lazy.

Get the Bits

As usual, the full code for this example is avaialble online here.

Why this matters

The historical coupling between Spring Batch and a relational database was always a pragmatic compromise, not a design ideal. The framework needs somewhere durable to remember what it did, and SQL was the path of least resistance. With the JobRepository abstraction now properly decoupled — and Spring Boot 4.1 shipping first-class autoconfiguration for MongoDB — teams running on document stores no longer have to keep a JDBC database around just to humor the batch tier.

Pick the database that fits your data. Spring Batch will fit itself around your choice.