Is your feature request related to a problem? Please describe.
In order to provide better protections against aggressive bots, DSpace should consider adding a "trusted IPs" configuration, to allow trusted bots easier harvesting access to DSpace (for SEO or similar) and also limit SSR to only "on campus" or other similar IP ranges.
What special behaviors would trusted bots possibly have?
- Trusted bots/IPs would always have SSR enabled (to allow for harvesting). This would also allow sites to potentially disable SSR site-wide, blocking out any un-trusted bots which cannot parse Javascript. Other known IPs (like "on campus" users) would also have SSR enabled.
- (Optionally) Trusted bots/IPs might be treated differently by our
ui > rateLimiter. Or, perhaps there's just a separate configuration for trusted bots in the rateLimiter, to allow sites to potentially loosen rules for trusted bots and tighten rules elsewhere.
- Trusted bots would be known, verifiable bots which could also be configured to serve cached SSR content (see
cache settings) if necessary for performance.
How would trusted bots be identified?
Trusted bots would need to be identified by IP address/range. Identifying bots by user-agent is no longer recommended practice because it's trivial to fake.
DSpace could populate an initial list of trusted bots from trusted sources such as:
Individual DSpace sites should also have a configuration (or similar) which allows them to add/remove/update additional "trusted bots". This allows sites to give access to other trusted crawlers / harvesters, or maybe remove access if a previously-trusted bot becomes problematic.
How will this alleviate issues with aggressive bots?
By providing a way to identify the good/trusted bots, DSpace sites could then turn off SSR without any impact on SEO (because sites could configure SSR to only be used for trusted bots).
Turning off SSR helps to deter aggressive bots since many of those bots currently attempt to harvest/parse HTML. Even if they do parse Javascript, they'd do all that processing on the client-side, which would not overwhelm the node server where dspace-angular runs.
Describe the solution you'd like
While I don't have an exact design in mind, we may wish to consider if a list of "trusted IPs" is best kept on the backend or fronted.
* If kept on the frontend, it'd likely need to be a large JSON file which can be parsed at runtime (and rebuilt on demand from configs or similar).
* If kept on the backend, we'd have more flexibility of keeping it in the database or configuration (or a bit of both). However, we'd still need to pass the raw data to the frontend via JSON.
Rough Implementation Proposal
For simplicity, consider storing the list of trusted IP addresses as a JSON file, likely on the backend (as it will need a script to regularly update it, and most of our regular cronjob-style scripts are on the backend).
The JSON file we store locally should use the JSON-Based format for publishing IP Ranges of Automated HTTP Clients, as that's the same format used by all search engines.
- Create a new set of configurations in
dspace.cfg (or a similar *.cfg file) which is structured similar to this:
# Configure trusted IP sources, which are all defined below.
# Additional sources can be added by adding a new "webui.trusted-ipranges.source.*" config for each
webui.trusted-ipranges.sources = google, bing, duckduckgo, custom
# Google source is a JSON file which will be parsed
# (NOTE: All these JSON files are the same format for every search engine)
webui.trusted-ipranges.source.google = https://www.gstatic.com/ipranges/goog.json
# BingBot source
webui.trusted-ipranges.source.bing = https://www.bing.com/toolbox/bingbot.json
# DuckDuckGo source
webui.trusted-ipranges.source.duckduckgo = https://duckduckgo.com/duckduckbot.json
# Custom source is a list of IP addresses which may correspond to on-campus users or similar
webui.trusted-ipranges.source.custom = 192.168.0.0/24, 192.168.2.2
# Additional sources can be added by just adding new "webui.trusted-ipranges.source.*" configs
# The text after "source" allows you to label the meaning behind the IPs.
# webui.trusted-ipranges.source.custom2 = 192.168.5.5, 192.168.6.6
# Schedule for how frequently the trusted IP ranges should be regenerated via config
# This example shows it running at 2:15AM every Saturday
# When commented out, this feature would be disabled
# webui.trusted-ipranges.cron = 0 15 2 ? * SAT
- The DSpace backend needs a script (e.g.
./dspace update-trusted-ipranges) which can loop through the webui.trusted-ipranges.sources and gather ALL IP ranges together into a single, large JSON file (named something like [dspace]/config/trusted-ipranges.json).
- This script should have a "cron" setting (see example above), allowing it to be run on a semi-regular basis, e.g. about once per week, because the IPs of search engines don't change frequently. The script should overwrite the existing JSON file each time it runs.
- Think of this script as similar to the
generate-sitemaps script which auto-updates Sitemaps on a regular basis also using a "cron" configuration.
- The single combined JSON file would use the standard JSON format used by all search engines, and look similar to below. Notice how the "service" in this JSON format would list the source of the IP address/range. That can allow the UI to potentially have different configurations for different services (see examples below).
{
"creationTime": "2024-01-03T10:00:00.121331",
"prefixes": [
{
"ipv4Prefix": "157.55.39.0/24",
"service": "bing"
},
{
"ipv4Prefix": "8.8.8.0/24",
"service": "google"
},
{
"ipv4Prefix": "104.43.54.127/32",
"service": "duckduckgo"
},
{
"ipv4Prefix": "192.168.0.0/24",
"service": "custom"
},
...
}
- The backend will need to make the
trusted-ipranges.json available to the frontend, perhaps proxied similar to how sitemap files generated on the backend are available in the frontend.
- This might mean making this JSON file visible publicly. However, if we determine that is a security or privacy risk we likely could limit its access only to the server where the UI is running, as this file primarily needs to be visible to the
server.ts file in the UI.
- The User Interface will need to be updated to request this
trusted-ipranges.json and parse it to determine whether SSR is enabled. An additional SSR config will need to be added similar to this:
ssr:
# When this is enabled, only trusted IPs listed in `trusted-ipranges.json` will trigger SSR.
enableTrustedIPRanges: true
# (OPTIONAL) we could even make it configurable based on service type in "trusted-ipranges.json"
# This example would only enable SSR for "google" and "custom" services in the "trusted-ipranges.json" file.
enableTrustedIPRangeServices:
- service: "google"
- service: "custom"
- (OPTIONAL) We optionally could update the "cache" settings to customize which cache is used for which service.
cache:
serverSide:
# This example shows us configuring that the "google" and "bing" IPs should all use the "botCache"
botCache:
max: 1000
trustedIPRangeServices:
- service: "google"
- service: "bing"
# But, the "custom" IPs should use the "anonymousCache" instead
anonymousCache:
max: 100
trustedIPRangeServices:
- service: "custom"
- Finally, obviously in the UI's
server.ts, we'd need a script/library that can quickly determine if the current user IP matches any of those in the trusted-ipranges.json. If a match is found, then SSR is triggered (and caching is triggered, if enabled). If not, then it falls back to CSR only.
Describe alternatives or workarounds you've considered
DSpace doesn't currently have a solution that lines up with this idea:
- DSpace can support only enabling SSR for some pages. But, that's a global config and doesn't allow you to change the settings per IP
- DSpace can support rateLimiting as well, but again it's a global config. No way to have different settings for different IP ranges
- DSpace also supports basic cached of SSR with a bot-specific configuration. But, that bot-specific configuration uses the
isbot library which tries to identify bots based on user-agent (which is less reliable and may not work at all for highly aggressive bots)
Additional information
This issue is related to #4565 and comes out of discussions with the Google Scholar team on Sept 8, 2026.
Is your feature request related to a problem? Please describe.
In order to provide better protections against aggressive bots, DSpace should consider adding a "trusted IPs" configuration, to allow trusted bots easier harvesting access to DSpace (for SEO or similar) and also limit SSR to only "on campus" or other similar IP ranges.
What special behaviors would trusted bots possibly have?
ui > rateLimiter. Or, perhaps there's just a separate configuration for trusted bots in therateLimiter, to allow sites to potentially loosen rules for trusted bots and tighten rules elsewhere.cachesettings) if necessary for performance.How would trusted bots be identified?
Trusted bots would need to be identified by IP address/range. Identifying bots by user-agent is no longer recommended practice because it's trivial to fake.
DSpace could populate an initial list of trusted bots from trusted sources such as:
Individual DSpace sites should also have a configuration (or similar) which allows them to add/remove/update additional "trusted bots". This allows sites to give access to other trusted crawlers / harvesters, or maybe remove access if a previously-trusted bot becomes problematic.
How will this alleviate issues with aggressive bots?
By providing a way to identify the good/trusted bots, DSpace sites could then turn off SSR without any impact on SEO (because sites could configure SSR to only be used for trusted bots).
Turning off SSR helps to deter aggressive bots since many of those bots currently attempt to harvest/parse HTML. Even if they do parse Javascript, they'd do all that processing on the client-side, which would not overwhelm the node server where
dspace-angularruns.Describe the solution you'd like
While I don't have an exact design in mind, we may wish to consider if a list of "trusted IPs" is best kept on the backend or fronted.
* If kept on the frontend, it'd likely need to be a large JSON file which can be parsed at runtime (and rebuilt on demand from configs or similar).
* If kept on the backend, we'd have more flexibility of keeping it in the database or configuration (or a bit of both). However, we'd still need to pass the raw data to the frontend via JSON.
Rough Implementation Proposal
For simplicity, consider storing the list of trusted IP addresses as a JSON file, likely on the backend (as it will need a script to regularly update it, and most of our regular cronjob-style scripts are on the backend).
The JSON file we store locally should use the JSON-Based format for publishing IP Ranges of Automated HTTP Clients, as that's the same format used by all search engines.
dspace.cfg(or a similar*.cfgfile) which is structured similar to this:./dspace update-trusted-ipranges) which can loop through thewebui.trusted-ipranges.sourcesand gather ALL IP ranges together into a single, large JSON file (named something like[dspace]/config/trusted-ipranges.json).generate-sitemapsscript which auto-updates Sitemaps on a regular basis also using a "cron" configuration.trusted-ipranges.jsonavailable to the frontend, perhaps proxied similar to how sitemap files generated on the backend are available in the frontend.server.tsfile in the UI.trusted-ipranges.jsonand parse it to determine whether SSR is enabled. An additional SSR config will need to be added similar to this:server.ts, we'd need a script/library that can quickly determine if the current user IP matches any of those in thetrusted-ipranges.json. If a match is found, then SSR is triggered (and caching is triggered, if enabled). If not, then it falls back to CSR only.Describe alternatives or workarounds you've considered
DSpace doesn't currently have a solution that lines up with this idea:
isbotlibrary which tries to identify bots based onuser-agent(which is less reliable and may not work at all for highly aggressive bots)Additional information
This issue is related to #4565 and comes out of discussions with the Google Scholar team on Sept 8, 2026.