Skip to content

How to Configure DI-engine for Single Custom Environment Setup in Training and Evaluation? #865

Description

@Nnasan

Hello, I am currently working on a project using DI-engine and I have a question regarding training and evaluation with a single environment.

In the DI-engine framework, I understand that the collector_env_num and evaluator_env_num parameters are used to manage the number of environments during training and evaluation. However, in my case, I only have one environment instance available, and I am wondering if it is possible to run training and evaluation with just one environment?

I tried configuring collector_env_num and evaluator_env_num to 1, but I am not sure if there are any issues with resource contention when switching between training and evaluation modes, especially since both use the same environment instance.

Could you provide guidance on how to properly configure DI-engine for a single environment setup, and whether there are any specific steps or considerations I should be aware of?

Activity

  1. PaParaZz1 commented on Apr 12, 2025

    @PaParaZz1
    Member

    Configuring collector_env_num and evaluator_env_num to 1 may lead to conflicts for you environment. I think you should write a special entry file for your usage: just use 1 collector and don't use evaluator, then use the reset_policy method to control the concrete behaviors for collection or evaluation.

  2. added
    discussionDiscussion of a typical issue
    envQuestions about RL environment
    on Apr 12, 2025
  3. Nnasan commented on Apr 13, 2025

    @Nnasan
    Author

    Thank you for your reply! Could you please provide a more detailed explanation on how to handle this situation, or share any example code? That would be very helpful.

  4. PaParaZz1 commented on May 28, 2025

    @PaParaZz1
    Member

    @Nnasan Here is a naive demo code:

    shared_env = create_env_manager(...)
    collector = create_serial_collector(env=shared_env, ...)
    evaluator = create_serial_evaluator(env=shared_env, ...)
    
    while True:
        collect_kwargs = commander.step()
        # Evaluate policy performance
        evaluator.reset_policy()  # call policy's `_reset_eval` method
        if evaluator.should_eval(learner.train_iter):
            stop, eval_info = evaluator.eval(learner.save_checkpoint, learner.train_iter, collector.envstep)
            if stop:
                break
        # Collect data by default config n_sample/n_episode
        collector.reset_policy()  # call policy's `_reset_collect` method
        new_data = collector.collect(train_iter=learner.train_iter, policy_kwargs=collect_kwargs)
        replay_buffer.push(new_data, cur_collector_envstep=collector.envstep)
        # Learn policy from collected data
        for i in range(cfg.policy.learn.update_per_collect):
            # Learner will train ``update_per_collect`` times in one iteration.
            train_data = replay_buffer.sample(learner.policy.get_attribute('batch_size'), learner.train_iter)
            learner.train(train_data, collector.envstep)
  5. Nnasan commented on Jun 13, 2025

    @Nnasan
    Author

    我按照你的方式,创建evaluator 时报错了
    packages\ding\envs\env_manager\subprocess_env_manager.py", line 211, in launch
    assert self._closed, "please first close the env manager"
    AssertionError: please first close the env manager
    我问了下ai 这似乎是说构造 evaluator 时传入了一个 已经被用作 collector 的环境对象 shared_env,但 evaluator 又尝试重新 launch(启动)它。
    此外还有个小小的问题: reset_policy此时是切换策略吗,还是重置策略?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    discussionDiscussion of a typical issueenvQuestions about RL environment

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions