Skip to content

Consider creating a trusted list of IPs for fine tuned SSR (to filter out aggressive bots) #6212

Description

@tdonohue

Is your feature request related to a problem? Please describe.

In order to provide better protections against aggressive bots, DSpace should consider adding a "trusted IPs" configuration, to allow trusted bots easier harvesting access to DSpace (for SEO or similar) and also limit SSR to only "on campus" or other similar IP ranges.

What special behaviors would trusted bots possibly have?

  1. Trusted bots/IPs would always have SSR enabled (to allow for harvesting). This would also allow sites to potentially disable SSR site-wide, blocking out any un-trusted bots which cannot parse Javascript. Other known IPs (like "on campus" users) would also have SSR enabled.
  2. (Optionally) Trusted bots/IPs might be treated differently by our ui > rateLimiter. Or, perhaps there's just a separate configuration for trusted bots in the rateLimiter, to allow sites to potentially loosen rules for trusted bots and tighten rules elsewhere.
  3. Trusted bots would be known, verifiable bots which could also be configured to serve cached SSR content (see cache settings) if necessary for performance.

How would trusted bots be identified?

Trusted bots would need to be identified by IP address/range. Identifying bots by user-agent is no longer recommended practice because it's trivial to fake.

DSpace could populate an initial list of trusted bots from trusted sources such as:

Individual DSpace sites should also have a configuration (or similar) which allows them to add/remove/update additional "trusted bots". This allows sites to give access to other trusted crawlers / harvesters, or maybe remove access if a previously-trusted bot becomes problematic.

How will this alleviate issues with aggressive bots?

By providing a way to identify the good/trusted bots, DSpace sites could then turn off SSR without any impact on SEO (because sites could configure SSR to only be used for trusted bots).

Turning off SSR helps to deter aggressive bots since many of those bots currently attempt to harvest/parse HTML. Even if they do parse Javascript, they'd do all that processing on the client-side, which would not overwhelm the node server where dspace-angular runs.

Describe the solution you'd like

While I don't have an exact design in mind, we may wish to consider if a list of "trusted IPs" is best kept on the backend or fronted.
* If kept on the frontend, it'd likely need to be a large JSON file which can be parsed at runtime (and rebuilt on demand from configs or similar).
* If kept on the backend, we'd have more flexibility of keeping it in the database or configuration (or a bit of both). However, we'd still need to pass the raw data to the frontend via JSON.

Rough Implementation Proposal

For simplicity, consider storing the list of trusted IP addresses as a JSON file, likely on the backend (as it will need a script to regularly update it, and most of our regular cronjob-style scripts are on the backend).

The JSON file we store locally should use the JSON-Based format for publishing IP Ranges of Automated HTTP Clients, as that's the same format used by all search engines.

  1. Create a new set of configurations in dspace.cfg (or a similar *.cfg file) which is structured similar to this:
    # Configure trusted IP sources, which are all defined below.
    # Additional sources can be added by adding a new "webui.trusted-ipranges.source.*" config for each
    webui.trusted-ipranges.sources = google, bing, duckduckgo, custom
    
    # Google source is a JSON file which will be parsed
    # (NOTE: All these JSON files are the same format for every search engine)
    webui.trusted-ipranges.source.google = https://www.gstatic.com/ipranges/goog.json
    
    # BingBot source
    webui.trusted-ipranges.source.bing = https://www.bing.com/toolbox/bingbot.json
    
    # DuckDuckGo source
    webui.trusted-ipranges.source.duckduckgo = https://duckduckgo.com/duckduckbot.json
    
    # Custom source is a list of IP addresses which may correspond to on-campus users or similar
    webui.trusted-ipranges.source.custom = 192.168.0.0/24, 192.168.2.2
    
    # Additional sources can be added by just adding new "webui.trusted-ipranges.source.*" configs
    # The text after "source" allows you to label the meaning behind the IPs.
    # webui.trusted-ipranges.source.custom2 = 192.168.5.5, 192.168.6.6
    
    # Schedule for how frequently the trusted IP ranges should be regenerated via config
    # This example shows it running at 2:15AM every Saturday
    # When commented out, this feature would be disabled
    # webui.trusted-ipranges.cron = 0 15 2 ? * SAT
    
  2. The DSpace backend needs a script (e.g. ./dspace update-trusted-ipranges) which can loop through the webui.trusted-ipranges.sources and gather ALL IP ranges together into a single, large JSON file (named something like [dspace]/config/trusted-ipranges.json).
    • This script should have a "cron" setting (see example above), allowing it to be run on a semi-regular basis, e.g. about once per week, because the IPs of search engines don't change frequently. The script should overwrite the existing JSON file each time it runs.
    • Think of this script as similar to the generate-sitemaps script which auto-updates Sitemaps on a regular basis also using a "cron" configuration.
  3. The single combined JSON file would use the standard JSON format used by all search engines, and look similar to below. Notice how the "service" in this JSON format would list the source of the IP address/range. That can allow the UI to potentially have different configurations for different services (see examples below).
    {
       "creationTime": "2024-01-03T10:00:00.121331",
       "prefixes": [
         {
           "ipv4Prefix": "157.55.39.0/24",
           "service": "bing"
         },
         {
           "ipv4Prefix": "8.8.8.0/24",
           "service": "google"
         },
         {
           "ipv4Prefix": "104.43.54.127/32",
           "service": "duckduckgo"
         },
         {
           "ipv4Prefix": "192.168.0.0/24",
           "service": "custom"
         },
         ...
    }
    
  4. The backend will need to make the trusted-ipranges.json available to the frontend, perhaps proxied similar to how sitemap files generated on the backend are available in the frontend.
    • This might mean making this JSON file visible publicly. However, if we determine that is a security or privacy risk we likely could limit its access only to the server where the UI is running, as this file primarily needs to be visible to the server.ts file in the UI.
  5. The User Interface will need to be updated to request this trusted-ipranges.json and parse it to determine whether SSR is enabled. An additional SSR config will need to be added similar to this:
    ssr: 
      # When this is enabled, only trusted IPs listed in `trusted-ipranges.json` will trigger SSR.
      enableTrustedIPRanges: true
      # (OPTIONAL) we could even make it configurable based on service type in "trusted-ipranges.json"
      # This example would only enable SSR for "google" and "custom" services in the "trusted-ipranges.json" file.
      enableTrustedIPRangeServices:
        - service: "google"
        - service: "custom"
    
  6. (OPTIONAL) We optionally could update the "cache" settings to customize which cache is used for which service.
    cache: 
      serverSide:
        # This example shows us configuring that the "google" and "bing" IPs should all use the "botCache"
        botCache:
          max: 1000
          trustedIPRangeServices:
            - service: "google"
            - service: "bing"
        # But, the "custom" IPs should use the "anonymousCache" instead
        anonymousCache:
          max: 100
          trustedIPRangeServices:
            - service: "custom"
    
  7. Finally, obviously in the UI's server.ts, we'd need a script/library that can quickly determine if the current user IP matches any of those in the trusted-ipranges.json. If a match is found, then SSR is triggered (and caching is triggered, if enabled). If not, then it falls back to CSR only.

Describe alternatives or workarounds you've considered

DSpace doesn't currently have a solution that lines up with this idea:

  • DSpace can support only enabling SSR for some pages. But, that's a global config and doesn't allow you to change the settings per IP
  • DSpace can support rateLimiting as well, but again it's a global config. No way to have different settings for different IP ranges
  • DSpace also supports basic cached of SSR with a bot-specific configuration. But, that bot-specific configuration uses the isbot library which tries to identify bots based on user-agent (which is less reliable and may not work at all for highly aggressive bots)

Additional information

This issue is related to #4565 and comes out of discussions with the Google Scholar team on Sept 8, 2026.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions