LangChain4j Ingest
Ingest documents into a LangChain4j EmbeddingStore: split, embed and store the message body.
What’s inside
-
LangChain4j Ingest component, URI syntax:
langchain4j-ingest:pipelineName
Please refer to the above links for usage and configuration details.
Maven coordinates
<dependency>
<groupId>org.apache.camel.springboot</groupId>
<artifactId>camel-langchain4j-ingest-starter</artifactId>
</dependency> Spring Boot Auto-Configuration
The starter supports 19 options, which are listed below.
| Name | Description | Default | Type |
|---|---|---|---|
camel.component.langchain4j-ingest.autowired-enabled | Whether autowiring is enabled. This is used for automatic autowiring options (the option must be marked as autowired) by looking up in the registry to find if there is a single instance of matching type, which then gets configured on the component. This can be used for automatic configuring JDBC data sources, JMS connection factories, AWS Clients, etc. | true | Boolean |
camel.component.langchain4j-ingest.configuration | The configuration. The option is a org.apache.camel.component.langchain4j.ingest.LangChain4jIngestConfiguration type. | LangChain4jIngestConfiguration | |
camel.component.langchain4j-ingest.content-type | MIME type of a media body, such as audio/wav or image/png, handed to the embedding model; matched case-insensitively, parameters after a semicolon are dropped. When not set, it is derived from the document id’s file extension through Camel’s own MIME table (the x- audio variants rewritten to their registered form): wav, mp3, flac, ogg, opus, m4a, aac, aiff, png, jpg, gif, webp, mp4, mov, webm (video), pdf and every other extension the table knows. A document whose type cannot be determined, whose type is no medium (text/plain, say), or whose medium the model does not declare, fails the exchange before its dedup claim and before its body is read. Only valid with modality=media - with modality=text the endpoint refuses to start, the option being a sign that modality=media was forgotten. | String | |
camel.component.langchain4j-ingest.document-filter | A Predicate deciding whether a delivery is ingested, referenced as #bean:name and evaluated with the message body available. A rejected delivery is answered with a filtered result and releases its dedup claim. Runs after the id patterns and after the dedup claim, so a duplicate is answered skipped without the filter being evaluated. The body is still as the consumer delivered it, a file or stream say, not yet read as text or bytes. Not looked up by type on purpose - an application may hold unrelated predicates. The option is a org.apache.camel.Predicate type. | Predicate | |
camel.component.langchain4j-ingest.document-id-header | Name of the header carrying the document id, such as CamelAwsS3Key for an S3 consumer or CamelKafkaKey for a Kafka one. The CamelLangChain4jIngestDocumentId exchange property, when set, takes precedence - a route that parses documents captures the id into that property before the parse, so a document cannot forge its own identity. An exchange without an id fails. | CamelLangChain4jIngestDocumentId | String |
camel.component.langchain4j-ingest.document-splitter | The DocumentSplitter deciding how a document becomes segments, referenced as #bean:name - LangChain4j ships alternatives beside the default recursive one. When set, maxSegmentSize and maxOverlapSize are ignored (they parameterize the default splitter only). Segments returned without the identity metadata are re-stamped, so a custom splitter cannot break citation. Not looked up by type on purpose - an application may hold unrelated splitters. The option is a dev.langchain4j.data.document.DocumentSplitter type. | DocumentSplitter | |
camel.component.langchain4j-ingest.embedding-batch-size | How many segments are embedded per request to the embedding model. Providers with generous per-request limits ingest large documents faster with a bigger batch; a batch carries at most embeddingBatchSize x maxSegmentSize characters, so tune the two together against the provider’s token limits. | 32 | Integer |
camel.component.langchain4j-ingest.embedding-model | The EmbeddingModel to embed segments with. When not set, the single bean of that type in the registry is used; zero or several beans fail the endpoint start with an error naming this option. The option is a dev.langchain4j.model.embedding.EmbeddingModel type. | EmbeddingModel | |
camel.component.langchain4j-ingest.embedding-store | The EmbeddingStore to write segments to. When not set, the single bean of that type in the registry is used; zero or several beans fail the endpoint start with an error naming this option. The option is a dev.langchain4j.store.embedding.EmbeddingStore<dev.langchain4j.data.segment.TextSegment> type. | TextSegment> | |
camel.component.langchain4j-ingest.enabled | Whether to enable auto configuration of the langchain4j-ingest component. This is enabled by default. | Boolean | |
camel.component.langchain4j-ingest.exclude-id | Comma-separated list of Ant-style patterns for document ids to skip. Exclusion wins over includeId. A matching delivery is answered with a filtered result, before the dedup claim and without reading the body. Matching is case-sensitive, and a pattern only matches an id that agrees with it on a leading path separator. See the component documentation for pattern examples. | String | |
camel.component.langchain4j-ingest.idempotent-repository | The IdempotentRepository remembering already ingested document ids, referenced as #bean:name. When set, a delivery whose id was already written is answered with a skipped result instead of being re-ingested: first write wins per id. A blank document releases its claim, so a later, populated delivery under the same id still ingests. The claim is eager: a duplicate racing an in-flight first delivery is answered skipped even if that delivery then fails - with an at-least-once source the skipped duplicate is acknowledged and the failed original may be the only other copy, so pair eager deduplication with a source that redelivers on failure. Not looked up by type on purpose - an application may hold unrelated idempotent repositories. The repository is started but never stopped by the endpoint (it may be shared); a persistent repository’s lifecycle belongs to whoever created it. The option is a org.apache.camel.spi.IdempotentRepository type. | IdempotentRepository | |
camel.component.langchain4j-ingest.include-id | Comma-separated list of Ant-style patterns the document id must match to be ingested. A non-matching delivery is answered with a filtered result, before the dedup claim and without reading the body. When not set, every id is accepted. Matching is case-sensitive, and a pattern only matches an id that agrees with it on a leading path separator - an id derived from an absolute path needs a pattern starting with one. See the component documentation for pattern examples. | String | |
camel.component.langchain4j-ingest.lazy-start-producer | Whether the producer should be started lazy (on the first message). By starting lazy you can use this to allow CamelContext and routes to startup in situations where a producer may otherwise fail during starting and cause the route to fail being started. By deferring this startup to be lazy then the startup failure can be handled during routing messages via Camel’s routing error handlers. Beware that when the first message is processed then creating and starting the producer may take a little time and prolong the total processing time of the processing. | false | Boolean |
camel.component.langchain4j-ingest.max-document-size | Maximum size of one document: characters of the text about to be split, or bytes of a media body with modality=media; 0, the default, means no limit. The pipeline holds a document in memory whole, so the cap is the protection against oversized - on a consumer-fed pipeline, attacker-sized - payloads. An oversized document fails the exchange cleanly and, with a repository configured, releases its dedup claim. With modality=media the size a file consumer announces in CamelFileLength is checked before the body is read. | 0 | Integer |
camel.component.langchain4j-ingest.max-overlap-size | How much of the previous segment each segment repeats, in characters. Overlap keeps a sentence split across a boundary retrievable from either side. | 50 | Integer |
camel.component.langchain4j-ingest.max-segment-size | Maximum size of one segment, in characters. | 500 | Integer |
camel.component.langchain4j-ingest.min-document-size | Minimum size of one document in characters (bytes with modality=media); 0, the default, means no minimum. A shorter document - boilerplate too small to carry retrievable content - is answered with a filtered result instead of being written, and releases its dedup claim like a blank one. | 0 | Integer |
camel.component.langchain4j-ingest.modality | What the message body is. text, the default, is read as a String, split into segments and embedded segment by segment. media is read as bytes and embedded whole, as one vector: audio, an image, video or a PDF, told apart by the MIME type, each needing an embedding model whose supportedContentTypes() include the matching type - the endpoint refuses to start with a text-only model. In media mode the splitter options and embeddingBatchSize do not apply, documentSplitter must not be set, and maxDocumentSize and minDocumentSize count bytes. | text | IngestModality |