Most search projects stall at the same point. The engine is fine. The problem is that the content lives in six places: a product database, a help center, a CMS, a ticketing tool, a shared drive and a marketing site. Nobody wants to write and babysit six pipelines. Search connectors exist to make that part boring.
What a search connector actually does
A connector is a small, source-specific program with one job: read records from a system and hand them to your search index in a shape the index understands. That sounds simple, and the reading part usually is. The work is in everything around it.
Take a Postgres table of orders. The first sync reads every row. The second sync should not. A good connector keeps a cursor, a timestamp or a row version, and only picks up what changed. Without that, a nightly full reload of a large table becomes the most expensive thing in your stack.
Then there is shape. Source records rarely match what search wants. A price stored as a string needs to become a number so you can filter on it. A nested author object needs to be flattened or kept as a nested field on purpose. The connector should discover the source schema first, show it to you, and let you confirm the mapping before a single record is written.
Pull, push, or both
There are two honest ways to keep an index current. Scheduled pull means the connector wakes up on a timer and asks the source what changed. It works with any source that has an API or a database you can query, and it is easy to reason about. Webhook push means the source calls you the moment something changes. It is faster, but only some systems offer it, and you still need a periodic pull to catch anything the webhook missed.
The right answer is usually both: push for freshness, pull for correctness. A connector that supports only one of them will eventually drift from the source, and you will find out from a customer.
Where connectors go wrong
Three failure modes show up again and again.
Credentials attached to one person. When the engineer who authorized the integration leaves, the sync stops. Authorization should belong to the organization, and one authorized connection should be reusable across several connector instances.
Schema changes that block writes. A new column at the source should not take search down. The safe pattern is to write each version of the schema into its own collection and point an alias at the current one. Reindex into the new version, flip the alias, and readers never notice.
Silent partial syncs. If a run indexes 40,000 of 42,000 records and reports success, that is a bug you will chase for weeks. Every run needs a record of what it read, what it wrote, what failed and why.
How this works in Lunexa today
For public websites, Smart Crawler is the connector. Give it a URL, let it sample pages and propose a schema, confirm the fields, and schedule re-crawls. For systems you control, the bulk JSONL import and the documents API take create, upsert and update actions, so a short script against your database or CMS export is enough to keep a collection current. Schema import lets you paste a collection definition instead of clicking through field by field.
Underneath, the same primitives apply to every source: collections scoped to a project, collection aliases for zero-downtime reindexing, and scoped search keys so the browser never sees an admin credential. Any connector for a third-party source should follow the same model: organization-level authorization, one instance per source resource, a schema you confirm before the first sync, and a versioned collection behind an alias.
If you are deciding how to connect data sources to search, start with the sources that change most often and index those first. Everything else can be a weekly pull. See plans for crawl and record limits.