Please use this identifier to cite or link to this item: http://localhost:8081/jspui/handle/123456789/21313
Title: Online Learning for Noisy Polynomial
Authors: Chauhan, Deepanshu
Issue Date: Jun-2023
Publisher: IIT Roorkee
Abstract: We study deep contextual bandits, a class of contextual bandits where each context-arm pair is associated with a feature vector having an unknown reward-generating function. We compared already available algorithms like LinUCB, Neural UCB, Neural LinUCB, and Neural Linear. Implemented Neural LinUCB algorithm on the real-world dataset, which has finite action, and modified the algorithm for infinite action cases to optimize an unknown polynomial function. Our algorithm does not use Deep Neural Network to optimize the unknown reward function, as in the case of Neural LinUCB; instead, it relies on the reinforcement learning algorithm only. The objective of algorithm is to learn the maxima of the reward generating function which is a polynomial but degree of polynomial is unknown to the agent. The proposed algorithm is able to learn reward generating function in most of the cases.
URI: http://localhost:8081/jspui/handle/123456789/21313
Research Supervisor/ Guide: Gupta, Manu
metadata.dc.type: Dissertations
Appears in Collections:MASTERS' THESES (MFSDS & AI)

Files in This Item:
File Description SizeFormat 
21566005_Deepanshu Chauhan.pdf1.56 MBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.